Papers with Text Classification

300 papers
Locale-agnostic Universal Domain Classification Model in Spoken Language Understanding (N19-2)

Copied to clipboard

Challenge: Existing approaches to leveraging data across locales to improve domain classification accuracy are ineffective.
Approach: They propose a locale-agnostic universal domain classification model that leverages available data across locales sharing the same language to improve domain classification accuracy.
Outcome: The proposed model outperforms baseline models especially when classifying locale-specific domains and low-resourced domains.
Demo Application for the AutoGOAL Framework (2020.coling-demos)

Copied to clipboard

Challenge: AutoGOAL is a framework for automatically finding the best way to solve a given computational task.
Approach: They present a web demo that showcases the main characteristics of the AutoGOAL framework in Python and a graph-based representation for machine learning pipelines.
Outcome: The proposed framework can be applied to Natural Language Processing and structured classification problems.
Is Something Better than Nothing? Automatically Predicting Stance-based Arguments Using Deep Learning and Small Labelled Dataset (N18-2)

Copied to clipboard

Challenge: Argument mining is a subset of NLP that deals with extracting arguments from user-based content.
Approach: They propose to use weakly supervised and semi-supervised methods to automatically annotate reviews and provide large annotated datasets.
Outcome: The proposed methods can be used to learn better models for implicit/explicit opinion classification.
Towards Code-switched Classification Exploiting Constituent Language Resources (2020.aacl-srw)

Copied to clipboard

Challenge: Code-switching is a communicative phenomenon denoting a shift from one language to another within the same speech exchange.
Approach: They propose to convert code-switched data into its constituent high resource languages for use in both monolingual and cross-lingual settings.
Outcome: The proposed code-switching language can be used for multiple downstream tasks . the proposed language increases the F1 score by 22% and 42.5% compared to the state-of-the-art.
Token Sequence Labeling vs. Clause Classification for English Emotion Stimulus Detection (2020.starsem-1)

Copied to clipboard

Challenge: Emotion stimulus detection is the task of finding the cause of an emotion in a textual description.
Approach: They propose to evaluate whether clause classification or token sequence labeling is better for emotion stimulus detection in English.
Outcome: The proposed framework compares clause classification and token sequence labeling on four English datasets.
The Authors Matter: Understanding and Mitigating Implicit Bias in Deep Text Classification (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies on text classification have focused on the bias towards the individuals mentioned in the text content.
Approach: They propose a framework to mitigate implicit bias in text classification models based on demographic attributes of authors . they propose to use this framework to train deep text classifiers to make predictions on the right features .
Outcome: The proposed framework outperforms existing models significantly in fairness and performance.
Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning.
Approach: They propose a framework that combines the scalability of LLM-generated labels with the precision of human annotations to achieve higher speed and accuracy comparable to larger models.
Outcome: The proposed framework significantly improves accuracy across utterance-level dialogue tasks, including sentiment detection (over 2%), dialogue act classification (over 1.5%), etc.
Using Classifier Features to Determine Language Transfer on Morphemes (N18-4)

Copied to clipboard

Challenge: Using native English data, we identify an English learner’s native language background based solely on the learner's English writing samples.
Approach: They perform a Native Language Identification task where they identify an English learner’s native language background based only on the learner's English writing samples.
Outcome: The proposed task is connected to a position in second language acquisition research that holds all learners acquire English grammatical morphemes in the same order, regardless of native language background.
Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia (2022.naacl-main)

Copied to clipboard

Challenge: Using neural networks, we can model the impact of speaker role on language use through the game of Mafia.
Approach: They analyze the effect of speaker role on language use through the game of Mafia, in which players are assigned either an honest or a deceptive role.
Outcome: The proposed model outperforms a standard BERT-based text classification approach on two auxiliary tasks and identifies features that distinguish between player roles.
ToVo: Toxicity Taxonomy via Voting (2025.findings-naacl)

Copied to clipboard

Challenge: Existing toxic content detection models face limitations due to the closed-source nature of training data and the paucity of explanations for their evaluation mechanism.
Approach: They propose a mechanism that integrates voting and chain-of-thought processes to produce a high-quality open-source dataset for toxic content detection.
Outcome: The proposed model improves transparency and customizability while facilitating better fine-tuning for specific use cases.
Long-Distance Dependencies Don’t Have to Be Long: Simplifying through Provably (Approximately) Optimal Permutations (P19-2)

Copied to clipboard

Challenge: Neural models at the sentence level often need to model the interaction between words . however, there is no guarantee that the standard ordering of words is computationally efficient or optimal .
Approach: They propose to use a dependency parse as a proxy for inter-word dependencies in a sentence to simplify the sentence with combinatorial objectives imposed on the sentence-parse pair.
Outcome: The proposed model improves classification accuracy and reduces classification error by 2.0% over the previous state of the art.
Identifying Emotional and Polar Concepts via Synset Translation (2024.starsem-1)

Copied to clipboard

Challenge: Emotion identification and polarity classification seek to determine sentiment expressed by a writer.
Approach: They propose a translation-based method for labeling each individual word sense and lexical concept into 20 different languages and translate them into multilingual sentiment lexicons.
Outcome: The proposed method outperforms existing methods and is available on GitHub . it contains 12,429 emotional synsets and 15,567 polar synset.
Fine-Tuned Neural Models for Propaganda Detection at the Sentence and Fragment levels (D19-50)

Copied to clipboard

Challenge: The system was evaluated on a unified development set without distributing the gold labels.
Approach: They propose to use fine-grained propaganda detection to build models that can explain why an article is propagandistic.
Outcome: The proposed model performed on all eighteen propaganda techniques in the corpus of the shared task.
Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for NLP calibration ignore inherent human label variation (HLV) split votes are a problem in high-stakes domains such as legal and medical decisions .
Approach: They present a case outcome classification dataset with judges' vote distributions and build a taxonomy of disagreement with SV-specific subcategories.
Outcome: The proposed model is compared against a judge vote distribution and assesses the alignment of perceived difficulty between models and humans.
Pixie: Preference in Implicit and Explicit Comparisons (2022.acl-short)

Copied to clipboard

Challenge: Existing studies on preference classification focus on explicit comparisons, but recent work has focused on indirect comparisons which lack comparative quantifiers and adjectives.
Approach: They propose a manual annotated dataset for preference classification that includes 8,890 app reviews.
Outcome: The proposed model outperforms the state-of-the-art model and achieves a weighted average F1 score of 83.34%.
Predicting the Outcome of Deliberative Democracy: A Research Proposal (P19-2)

Copied to clipboard

Challenge: Deliberative dialogue is a structured, face-to-face method of public interaction that is fundamental to the concept of deliberative democracy.
Approach: They propose to use a combination of lexical, sentiment, durational and further ‘derivative’ features of adjacency pairs to train traditional classification models.
Outcome: The proposed method improves the accuracy of classification models and prediction tasks and shows that the task of recognising agreement is demanding but possible.
Improving Implicit Discourse Relation Classification by Modeling Inter-dependencies of Discourse Units in a Paragraph (N18-1)

Copied to clipboard

Challenge: Existing methods for predicting implicit discourse relations ignore wider paragraph contexts beyond the two discourse units examined for a discourse relation prediction.
Approach: They propose a paragraph-level neural network that models inter-dependencies between discourse units and discourse relation continuity and patterns and predicts a sequence of discourse relations in a sentence.
Outcome: The proposed model outperforms state-of-the-art systems on the benchmark corpus of PDTB.
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication (2025.coling-main)

Copied to clipboard

Challenge: Existing research indicates that disfluencies can constitute up to 5.9% of words in spontaneous speech, with repetitions accounting for over half of these disfluency.
Approach: They propose to use a dataset to analyze reduplication and repetition in speech using computational linguistics to evaluate transformer-based models.
Outcome: The proposed models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification.
Monitoring Hate Speech in Indonesia: An NLP-based Classification of Social Media Texts (2024.emnlp-demo)

Copied to clipboard

Challenge: a lack of mechanisms to track the spread and severity of hate speech complicates the formulation of effective solutions.
Approach: They have developed a universally robust hate speech classifier tailored for a narrower subset of texts that target vulnerable groups that have historically been the targets of hate speech in Indonesia.
Outcome: The proposed tool has persuaded the General Election Supervisory Body in Indonesia (BAWASLU) to collaborate with the Alliance of Independent Journalists (AJI) to monitor hate speech in vulnerable areas in the country known for hate speech dissemination or hate-related violence in the upcoming Indonesian regional elections.
Challenges and Remedies of Domain-Specific Classifiers as LLM Guardrails: Self-Harm as a Case Study (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have impressive capabilities in generating human-like text, but they pose significant risks in many domains and require guardrails throughout the lifecycle.
Approach: They propose to use a self-harm detector to test the performance of LLM guardrails in real-world environments.
Outcome: The proposed model performs poorly in open and closed domains and is almost unusable in the real world.
Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours (2022.emnlp-demos)

Copied to clipboard

Challenge: Label Sleuth is an open source system for labeling and creating text classifiers which does not require coding skills nor machine learning knowledge.
Approach: *Label Sleuth* is an open source system for labeling and creating text classifiers which does not require coding skills nor machine learning knowledge.
Outcome: *Label Sleuth* is an open source system for labeling and creating text classifiers.
Ukrainian Resilience: A Dataset for Detection of Help-Seeking Signals Amidst the Chaos of War (2024.findings-emnlp)

Copied to clipboard

Challenge: a new dataset is being developed to categorize posts that show distress or urgency . the dataset could improve humanitarian efforts, allowing for quicker and more targeted help .
Approach: They propose a dataset that brings together social media posts in the Ukrainian language for the detection of help-seeking posts in times of war.
Outcome: The proposed dataset can be used to improve humanitarian efforts . it can be compared with existing datasets and achieve an accuracy of 81.15% .
Multimodal Fallacy Classification in Political Debates (2024.eacl-short)

Copied to clipboard

Challenge: Recent advances in NLP suggest that some tasks, such as argument detection and relation classification, are better framed in a multimodal perspective.
Approach: They propose to use multimodal argument mining to capture paralinguistic aspects of fallacious arguments.
Outcome: The proposed multimodal argument mining improves argument detection and relation classification in political debates.
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification (2024.eacl-long)

Copied to clipboard

Challenge: Few-shot text classification systems are infeasible to deploy and use reliably due to their dependence on prompting and billion-parameter language models.
Approach: They propose a modification to SetFit that fine-tunes a Sentence Transformer under a contrastive learning paradigm and achieves similar results to more unwieldy systems.
Outcome: The proposed model fine-tunes a Sentence Transformer under a contrastive learning paradigm and achieves similar results to more unwieldy systems.
EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts (2022.aacl-short)

Copied to clipboard

Challenge: EmoNoBa is a dataset for fine-grained emotion detection on Bangla text . it is based on 22698 comments from social media sites on 12 domains .
Approach: They propose a manually annotated dataset of 22,698 Bangla comments from social media sites on 12 different domains to use for fine-grained emotion detection.
Outcome: The proposed dataset of 22,698 public comments on 12 domains shows that hand-crafted features perform better than neural networks and pre-trained language models.
Speak up, Fight Back! Detection of Social Media Disclosures of Sexual Harassment (N19-3)

Copied to clipboard

Challenge: #MeToo movement provides platform to narrate personal experiences of sexual harassment.
Approach: They propose a three-part ULMFiT architecture to tackle text subtleties in a classification task . they propose to annotate a manually annotated real-world dataset to test their approach .
Outcome: The proposed model outperforms existing models that rely on handcrafted stylistic features and is more accurate than generic models.
Label Noise in Context (2020.acl-demos)

Copied to clipboard

Challenge: Label noise—incorrectly or ambiguously labeled training examples—can negatively impact model performance.
Approach: They propose a noise-detection method that uses an example's neighborhood within the training set to reduce false positives and provide an explanation as to why the ex ample was flagged as noise.
Outcome: The proposed method outperforms the state-of-the-art on precision and F0.5-score on short-text classification datasets.
Multi-Dimensional Gender Bias Classification (2020.emnlp-main)

Copied to clipboard

Challenge: a novel framework decomposes gender bias in text along several pragmatic and semantic dimensions . language is a primary means by which people communicate, express identities and categorize themselves . unwanted gender biases can affect downstream applications, leading to poor user experiences .
Approach: They propose a framework that decomposes gender bias in text along several dimensions . they annotate eight large scale datasets with gender information and collect a benchmark .
Outcome: The proposed framework decomposes gender bias in text along several pragmatic and semantic dimensions.
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)

Copied to clipboard

Challenge: Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization.
Approach: They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool.
Outcome: The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application.
Content-Based Conflict of Interest Detection on Wikipedia (L18-1)

Copied to clipboard

Challenge: Conflict-of-Interest (CoI) editing is a problem on Wikipedia that is highly subjective . a key feature of Wiki sites is to allow people from all over the world to add or modify articles anonymously and without consequence.
Approach: They frame CoI detection as a binary classification problem and explore features for it . they find that stylometric features outperform other types of features and give an F-measure of 0.63 .
Outcome: The proposed method outperforms other features and gives an F-measure of 0.63 . the proposed method is not certain that the set of non-CoI articles contains any CoI articles .
Quality Estimation for Partially Subjective Classification Tasks via Crowdsourcing (2020.lrec-1)

Copied to clipboard

Challenge: a common approach to quality estimation is to ask multiple reviewers to evaluate the same artifacts.
Approach: They propose a probabilistic model for subjective classification tasks that incorporates the qualities of artifacts as well as the abilities and biases of creators and reviewers as latent variables to be jointly inferred.
Outcome: The proposed model estimates the quality of speech more effectively than a vote aggregation, measured by correlation with a fine-grained classification by experts.
A Fast, Compact, Accurate Model for Language Identification of Codemixed Text (D18-1)

Copied to clipboard

Challenge: a feed-forward network can label codemixed and monolingual text in 100 languages and 100 language pairs.
Approach: They propose a feed-forward network that can provide a language code for every token in a sentence . they show that the model can label both codemixed and monolingual text in 100 languages .
Outcome: The proposed model outperforms previous multilingual approaches in terms of accuracy and speed.
WikiRef: Wikilinks as a route to recommending appropriate references for scientific Wikipedia pages (C18-1)

Copied to clipboard

Challenge: Existing methods to enhance Wikipedia's reference section are not effective.
Approach: They propose a two-step approach that leverages the wikilinks present in a scientific Wikipedia target page and recommends highly relevant references to be included in that target page appropriately and automatically borrowed from the reference section of the wikipedia links.
Outcome: The proposed approach achieves a notably good performance on two datasets from Computer Science and Physics.
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research shows that named entities influence PLMs in many applications.
Approach: They propose a method to quantify biases associated with named entities from various countries using Twitter data instead of templates or specific datasets.
Outcome: The proposed method shows positive biases related to the language spoken in a country across all classifiers.
Privacy-Preserving Graph Convolutional Networks for Text Classification (2022.lrec-1)

Copied to clipboard

Challenge: Graph convolutional networks (GCNs) are powerful for representation learning on documents that naturally occur as graphs, but sensitive personal information is prone to privacy leaks.
Approach: They propose a method that adapts differentially-private gradient-based training to GCNs and conduct experiments using two optimizers on five NLP datasets in two languages.
Outcome: The proposed method improves baseline privacy bounds by 2.7 while retaining competitive F1 scores while providing strong privacy guarantees.
Contextual Modulation for Relation-Level Metaphor Identification (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to identifying metaphors in text ignore context where metaphor occurs . existing approaches focus on word-level identification without explicitly modelling interaction between metaphor components .
Approach: They propose a method for identifying relation-level metaphoric expressions of certain grammatical relations based on contextual modulation.
Outcome: The proposed architecture achieves state-of-the-art results on benchmark datasets.
Towards Unsupervised Text Classification Leveraging Experts and Word Embeddings (P19-1)

Copied to clipboard

Challenge: a new method for text classification uses supervised and semisupervised approaches to classify documents into categories.
Approach: They propose an unsupervised method to classify documents into categories simply described by a label.
Outcome: The proposed method increases F1-score over relying on human expertise and language models on standard corpora.
Analysis of Hierarchical Multi-Content Text Classification Model on B-SHARP Dataset for Early Detection of Alzheimer’s Disease (2020.aacl-main)

Copied to clipboard

Challenge: Existing studies on the detection of Alzheimer's disease focus on the diagnosis of dementia instead.
Approach: They propose to use a dataset to develop NLP models for the detection of Mild Cognitive Impairment (MCI) MCI is a progressive neurodegenerative disorder associated with memory loss and declines in major brain functions including semantic and pragmatic levels of language processing.
Outcome: The proposed model performs best on 74.1% of the 3 topics studied.
Where do LLMs Encode the Knowledge to Assess the Ambiguity? (2025.coling-industry)

Copied to clipboard

Challenge: False sizing of large language models can generate unreliable responses .
Approach: They propose a method to train large language models without ambiguity labels .
Outcome: The proposed method detects ambiguous input prompts better than representations from the final layer.
VertAttack: Taking Advantage of Text Classifiers’ Horizontal Vision (2024.naacl-long)

Copied to clipboard

Challenge: Current SOTA classifiers are subject to problems like bias and are vulnerable to adversarial attacks.
Approach: They propose an attack to mimic a classifier's character based attack and thenrewrite those words vertically.
Outcome: The proposed attack can drop the accuracy of 4 different transformer models on 5 datasets and preserve meaning.
A Fully Hyperbolic Neural Model for Hierarchical Multi-Class Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models for fine-grained entity typing have a hierarchical structure . prior work has integrated only explicit hierarchic information by formulating a hierarchy-aware loss or by representing instances and labels in a joint Euclidean embedding space.
Approach: They propose a fully hyperbolic model for multi-class multi-label classification that performs all operations in hyperbolical space.
Outcome: The proposed model performs all operations in hyperbolic space on two challenging datasets and shows it is comparable to state-of-the-art methods on fine-grained classification with remarkable reduction of parameter size.
Towards Context-Based Violence Detection: A Korean Crime Dialogue Dataset (2024.findings-eacl)

Copied to clipboard

Challenge: Currently, there are three main branches of violence detection, including surveillance of potential threats in offline situation and automatic prevention of harmful media.
Approach: They propose to use the Korean Crime Dialogue Dataset to classify violence that occurs in offline scenarios.
Outcome: The proposed dataset shows that understanding varying relationships among interlocutors improves the performance of crime dialogue classification.
#NotAWhore! A Computational Linguistic Perspective of Rape Culture and Victimization on Social Media (2020.acl-srw)

Copied to clipboard

Challenge: Recent surge in online forums and movements supporting sexual assault survivors has led to the emergence of a ‘virtual bubble’ where survivors can recount their stories.
Approach: They propose a transfer-learning based method to identify victim blaming language on Twitter and a single step transfer-based classification method to classify it.
Outcome: The proposed method is compared with various deep learning and machine learning models on a manually annotated domain-specific dataset.
Classifying Argumentative Relations Using Logical Mechanisms and Argumentation Schemes (2021.tacl-1)

Copied to clipboard

Challenge: Recent studies have focused on training complex neural networks on labeled data.
Approach: They propose to use logical mechanisms to classify argumentative relations without training on labeled data.
Outcome: The proposed method classifies argumentative relations without training on labeled data significantly better than unsupervised baselines.
Many Faces of Feature Importance: Comparing Built-in and Post-hoc Feature Importance in Text Classification (D19-1)

Copied to clipboard

Challenge: Feature importance is commonly used to explain machine predictions . however, the consistency of feature importance via different methods remains understudied .
Approach: They compare feature importance from built-in mechanisms and post-hoc methods that approximate model behavior to find similarities between models.
Outcome: The proposed methods show that features from traditional models are more similar with each other than with deep learning models.
CAL-Log: Cost-Aware Active Learning with Logarithmic Cognitive Effort Modeling and Online Adaptation to Human Annotation Behavior (2026.acl-srw)

Copied to clipboard

Challenge: Standard uncertainty sampling assumes that annotating a 500-word document requires the same effort as a 50-word tweet, leading to suboptimal resource allocation when documents vary in length.
Approach: They propose a cost-aware AL variant using logarithmic cost modeling where C(x) is the predicted annotation time for document x and L(x), is its token length.
Outcome: Experiments on ten text classification benchmarks show a 3.3 speedup over BADGE and 3.9 over Entropy sampling to reach F1=0.80, with large effect sizes.
A Knowledge-Augmented Neural Network Model for Implicit Discourse Relation Classification (C18-1)

Copied to clipboard

Challenge: Existing studies on implicit discourse relation classification have shown success using feedforward networks and convolutional neural networks.
Approach: They propose to augment input text with external knowledge and context and adopt a neural network model that can effectively handle the augmented text.
Outcome: The proposed model outperforms existing models on implicit discourse relation classification.
Domain-Agnostic Neural Architecture for Class Incremental Continual Learning in Document Processing Platform (2023.acl-industry)

Copied to clipboard

Challenge: Recent methods with stochastic gradient learning struggle in streaming data setups and are restricted to specific domains.
Approach: They propose a fully differentiable architecture that enables the training of high-performance classifiers when examples from each class are presented separately.
Outcome: The proposed architecture achieves SOTA results without a memory buffer and clearly outperforms the reference methods.
CoFE: A New Dataset of Intra-Multilingual Multi-target Stance Classification from an Online European Participatory Democracy Platform (2022.aacl-short)

Copied to clipboard

Challenge: Stance Recognition is a useful tool for many real-life applications, from misinformation detection to poll verification.
Approach: They propose to use an online debating platform where users can submit proposals and comment over proposals or over other comments.
Outcome: The proposed dataset contains 4.2k proposals and 20k comments on various topics.
PartisanLens: A Multilingual Dataset of Hyperpartisan and Conspiratorial Immigration Narratives in European Media (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for detecting hyperpartisan narratives and PRCTs are limited . hyperpartisan content promotes extreme views through one-sided, emotional language .
Approach: They propose a multilingual dataset of 1617 hyperpartisan news headlines in Spanish, Italian, and Portuguese annotated in multiple political discourse aspects.
Outcome: The proposed dataset is the first multilingual dataset of 1617 hyperpartisan headlines in Spanish, Italian, and Portuguese.
Cross-Modal Cloze Task: A New Task to Brain-to-Word Decoding (2022.findings-acl)

Copied to clipboard

Challenge: Existing work on decoding language from non-invasive brain activity is limited due to noisy nature of brain recordings.
Approach: They propose a cross-modal cloze task to predict a word from a neural image . they use a pre-trained language model to leverage the pre-training language model .
Outcome: The proposed method outperforms baselines on 20 participants from two brain imaging datasets.
Lexical Features Are More Vulnerable, Syntactic Features Have More Predictive Power (D19-55)

Copied to clipboard

Challenge: Existing metrics to quantify lexical diversity have been proposed.
Approach: They propose to examine how generic language characteristics are impacted by text alterations.
Outcome: The proposed models show that lexical features are more sensitive to text modifications than syntactic ones.
Thesis Proposal: Measuring Prejudice at Scale (2026.eacl-srw)

Copied to clipboard

Challenge: Existing methods for NLP fail to confirm construct validity, limiting the validity of the model.
Approach: They propose to shift from categorical classification to comparative scaling of grounded constructs by using prompt optimization and distillation approaches.
Outcome: The proposed pipeline is scalable for moving from categorical classification to theoretically grounded comparative measurement.
LionGuard: A Contextualized Moderation Classifier to Tackle Localized Unsafe Content (2025.coling-industry)

Copied to clipboard

Challenge: Existing moderation classes that perform well on English may not perform well in Singlish.
Approach: They propose a Singapore-contextualized moderation classifier that can guard against unsafe LLM usage.
Outcome: The proposed class outperforms existing modes on Singlish data by 14% and up to 51%.
Fine-Grained Temporal Orientation and its Relationship with Psycho-Demographic Correlates (N18-1)

Copied to clipboard

Challenge: Temporal orientation refers to an individual’s tendency to connect to the psychological concepts of past, present or future and affects personality, motivation, emotion, decision making and stress coping processes.
Approach: They propose to use a minimally supervised method to classify tweets in one of three temporal categories, past, present, and future, and a deep bi-directional long-term memory (BLSTM) to measure correlation between sentiment view of temporal orientation and different psycho-demographic factors.
Outcome: The proposed method achieves 78.27% accuracy on a manually created test set.
Predicting Customer Satisfaction with Soft Labels for Ordinal Classification (2023.acl-industry)

Copied to clipboard

Challenge: a typical call center only responds to 8% of customers with a customer satisfaction survey . a predictive algorithm that infers CSAT on the 1-5 scale is needed to minimize this data sparsity and response bias.
Approach: They propose an algorithm that infers CSAT on 1-5 scale on inbound calls to the call center . they reframe the problem into a binary class and map it back to five classes .
Outcome: The proposed model is able to support keycustomer workflows with high accuracy overmillions of calls a month.
LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification (2025.findings-naacl)

Copied to clipboard

Challenge: a lack of large-scale annotated datasets hinders effective training of ML models . despite advances in semantic segmentation, challenges persist in distinguishing between closely related roles .
Approach: They propose a large annotated dataset for semantic segmentation of legal documents . they use a rhetorical role classification model to compare performance against other models .
Outcome: The largest annotated dataset for this task outperforms models relying on sentence-level features.
Target-Based Offensive Language Identification (2023.acl-short)

Copied to clipboard

Challenge: Popular social media annotation taxonomies focus on the post level and token-level annotations are not available.
Approach: They propose a new dataset for Target-based Offensive language identification that uses post-level and token-level annotations to identify offensive language on Twitter.
Outcome: The proposed taxonomy can be used to annotate offensive language on English Twitter posts.
Classification of Moral Foundations in Microblog Political Discourse (P18-1)

Copied to clipboard

Challenge: a recent study shows correlation between political ideologies and moral foundations expressed in text . a moral foundation theory suggests that there are five basic moral values which underlie human moral perspectives .
Approach: They propose to model the moral foundations of tweets by using an annotation framework . they propose to use policy frames to predict the morality of political tweets .
Outcome: The proposed model can predict moral foundations of political tweets, the authors show . their model can be used to predict political slogans and political ideologies, they say .
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
Estimating User Communication Styles for Spoken Dialogue Systems (2020.lrec-1)

Copied to clipboard

Challenge: a neural network estimation system for spoken dialogues can be used to estimate the communication style of a user's interaction, but this is rarely implemented in a live system.
Approach: They propose a neural network approach to estimate the communication style of spoken interaction, namely elaborateness and directness.
Outcome: The proposed method can estimate the elaborateness and directness of spoken interaction and improve the results with additional linguistic features.
The Internal State of an LLM Knows When It’s Lying (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown exceptional performance in various tasks, but one of their main drawbacks is generating inaccurate or false information with a confident tone.
Approach: They propose to train a classifier that outputs the probability that a statement is truthful based on the hidden layer activations of the LLM as it reads or generates the statement.
Outcome: The proposed classifier achieves an average of 71% to 83% accuracy labeling which sentences are true versus false, depending on the LLM base model.
CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Large-Scale Multi-Label Text Classification (LMTC) tasks with hierarchical label spaces include automatic assignment of ICD-9 codes to discharge summaries.
Approach: They propose a set of metrics for hierarchical evaluation using the depth of the ontology to evaluate the predictions of neural LMTC models.
Outcome: The proposed metrics compare with previous evaluations on prior art models for ICD-9 coding in MIMIC-III and propose further avenues of research involving the proposed representation.
Detecting LLM-Assisted Cheating on Open-Ended Writing Tasks on Language Proficiency Tests (2024.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used for open-ended writing tasks . however, there are limitations in detecting LLM-generated samples .
Approach: They propose a framework for training LLM-generated text detectors that can detect LLM generated samples after being copy-typed.
Outcome: The proposed model outperforms the transformer-based classifier on a high-stakes online English proficiency test.
Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification (2024.naacl-short)

Copied to clipboard

Challenge: Emotion detection in textual data is pivotal for developing empathetic human-computer interaction systems.
Approach: They propose a method for categorizing emotions from textual data that acknowledges similarities and distinctions of various emotions.
Outcome: The proposed method preserves high accuracy in emotion prediction and significantly reduces errors in misclassification cases.
Predicting accuracy on large datasets from smaller pilot data (P18-2)

Copied to clipboard

Challenge: obtaining training data is often the most difficult part of an NLP or ML project . obtaining data is important to estimate how much training data a system will require to achieve a target accuracy.
Approach: They propose a performance extrapolation task to evaluate extrapolations on larger training sets.
Outcome: The proposed method can predict accuracy on larger training datasets.
Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks (2024.findings-eacl)

Copied to clipboard

Challenge: Prior work on multimodal content classification has not addressed these challenges.
Approach: They propose to use two auxiliary tasks to fine-tune multimodal models to address hidden cross-modal semantics and weak image-text relationships when modeling text and images.
Outcome: The proposed model improves by up to 2.6 F1 score across five diverse social media datasets.
Sequential Span Classification with Neural Semi-Markov CRFs for Biomedical Abstracts (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for dividing biomedical abstracts into rhetorical segments assign a rhetorical label to each sentence while considering context in the abstract.
Approach: They propose to use Neural Semi-Markov Conditional Random Fields to assign a rhetorical label to a span that consists of continuous sentences.
Outcome: The proposed method achieved the best micro sentence-F1 score and the best macro span-F1.
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem (2025.findings-emnlp)

Copied to clipboard

Challenge: a new pipeline can be used to create corpora for over-looked languages .
Approach: We propose a new pipeline that can filter a single snapshot in twohours.
Outcome: The proposed pipeline can filter a single snapshot in twohours.
Linguistically Motivated Features for Classifying Shorter Text into Fiction and Non-Fiction Genre (2022.coling-1)

Copied to clipboard

Challenge: linguistically motivated features are used to classify paragraph-level text into fiction and non-fiction genres.
Approach: They deploy linguistically motivated features to classify paragraph-level text into fiction and non-fiction genres using a logistic regression model.
Outcome: The proposed model gives 15.56% accuracy jump over baseline model . the proposed model also transfers over to another dataset, Baby BNC corpus .
Improving Compositional Generalization in Classification Tasks via Structure Annotations (2021.acl-short)

Copied to clipboard

Challenge: Compositional generalization is the ability to generalize systematically to a new data distribution by combining known components.
Approach: They propose to convert a natural language sequence-to-sequence dataset into a classification dataset that requires compositional generalization.
Outcome: The proposed model can generalize compositionally by providing hints on the structure of the input.
“Why do I feel offended?” - Korean Dataset for Offensive Language Identification (2023.findings-eacl)

Copied to clipboard

Challenge: Existing methods for detecting offensive content rely on labeled datasets, but few consider low-resource languages with relatively less data available for training.
Approach: They propose to use Korean as a dataset for offensive language identification . they propose to perform abusive language detection and sentiment analysis to help identify offensive languages.
Outcome: The proposed datasets improve the performance of offensive language identification in Korean, while the existing methods are limited.
Transformation Networks for Target-Oriented Sentiment Classification (P18-1)

Copied to clipboard

Challenge: a new model for sentiment classification uses attention instead of attention to classify sentiment polarities over individual opinion targets.
Approach: They propose a model that uses a CNN layer to extract salient features from transformed word representations from a bi-directional RNN layer.
Outcome: The proposed model achieves state-of-the-art on a few benchmarks.
CHEER-Ekman: Fine-grained Embodied Emotion Classification (2025.acl-short)

Copied to clipboard

Challenge: Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied.
Approach: They propose to extend existing binary embodied emotion dataset with Ekman’s six basic emotion categories.
Outcome: The proposed dataset outperforms existing methods with large language models.
Measuring What Counts: The Case of Rumour Stance Classification (2020.aacl-main)

Copied to clipboard

Challenge: Numerous methods have been proposed to predict the stance of replies towards a given rumour, but their performance is not optimal for the four-class imbalanced task of rumor stance classification.
Approach: They propose to use a four-class problem to predict the stance of replies towards a given rumour to help identify the most informative minority classes.
Outcome: The proposed methods are robust to imbalanced data and score higher systems capable of recognising the two most informative minority classes (support and deny).
Self-Governing Neural Networks for On-Device Short Text Classification (D18-1)

Copied to clipboard

Challenge: Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud.
Approach: They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing.
Outcome: The proposed models perform better on dialog act classification tasks while maintaining high accuracy.
Exploring Label Hierarchy in a Generative Way for Hierarchical Text Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for hierarchical text classification are lacking in the field of natural language processing.
Approach: They propose a hierarchy-aware T5 model with path-adaptive attention mechanism to exploit hierarchical dependency across different levels.
Outcome: The proposed model outperforms state-of-the-art models especially in Macro-F1 and low Macro.
What Motivates You? Benchmarking Automatic Detection of Basic Needs from Short Posts (2021.acl-short)

Copied to clipboard

Challenge: According to the self-determination theory, the levels of satisfaction of three basic needs (competence, autonomy and relatedness) have implications on people’s everyday life and career.
Approach: They propose to model a task that automatically detects three basic needs on short posts in English and then apply them to a binary task.
Outcome: The proposed model achieves similar performance as a trained human annotator in the real-world.
Multilingual Multi-class Sentiment Classification Using Convolutional Neural Networks (L18-1)

Copied to clipboard

Challenge: a new language-independent model for sentiment analysis is proposed for social media . a sentiment dictionary cannot list all the possible ways people can express their opinions .
Approach: They propose a language-independent model for multi-class sentiment analysis using a neural network architecture.
Outcome: The proposed model does not rely on language-specific features such as ontologies, dictionaries, or morphological or syntactic pre-processing.
Parameter Selection: Why We Should Pay More Attention to It (2021.acl-short)

Copied to clipboard

Challenge: supervised learning is a challenging process due to the huge number of parameter combinations.
Approach: They present an example of parameter selection in supervised learning . authors use a set of frequently occurring labels without a parameter tuning . they say this illustrates the seriousness of parameter tuning in a supervised field .
Outcome: The proposed study shows that without adequate attention, the research progress can be uncertain or even illusive.
Self-Governing Neural Networks for On-Device Short Text Classification (D18-1)

Copied to clipboard

Challenge: Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud.
Approach: They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing.
Outcome: The proposed models perform better on dialog act classification tasks while maintaining high accuracy.
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments (2025.acl-industry)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) systems aim to evaluate the quality of candidate writing using computational methods.
Approach: They propose a model that assigns a confidence score to each automated score to ensure it meets high reliability standards.
Outcome: The proposed model achieves an F1 score of 0.97 and releases 47% of predicted scores with 100% CEFR agreement and 99% with at least 95% CEFR agreeance compared to the standalone model where all predicted scores are released.
Supervised Domain Enablement Attention for Personalized Domain Classification (D18-1)

Copied to clipboard

Challenge: Recent IPDAs cover more than several thousands of diverse domains including Alexa Skills, Google Actions, and Cortana Skills.
Approach: They propose a supervised enablement attention mechanism that utilizes sigmoid activation for the attention weighting and self-distillation to leverage the attention information of other enabled domains.
Outcome: The proposed approach improves domain classification performance on real-world domains.
Putting Context in Context: the Impact of Discussion Structure on Text Classification (2024.eacl-long)

Copied to clipboard

Challenge: Current text classification approaches focus on the content to be classified, but contextual information is neglected in many cases.
Approach: They propose to integrate contextual information into a transformer-based model by feeding it as natural language input into . they also experiment with different amounts of training data and analyse local discussion networks in a privacy-compliant way.
Outcome: The proposed model can be generalized to other datasets and is privacy-compliant.
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts.
Approach: They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues.
Outcome: The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks.
Classifying the Informative Behaviour of Emoji in Microblogs (L18-1)

Copied to clipboard

Challenge: Emoji are pictographs used in microblogs as emotion markers, but can also represent a wider range of concepts.
Approach: They analyze a corpus of tweets pairs and classify emoji with respect to redundancy . they propose to further investigate the informative behaviour of e-mails using eoji .
Outcome: The proposed model achieved an F-score of 0.7 for emoji use in 2475 tweets pairs.
CVAE-based Re-anchoring for Implicit Discourse Relation Classification (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that training implicit discourse relation classifiers suffers from data sparsity.
Approach: They propose a re-anchoring strategy to reduce the risk of erroneous sampling . they use Conditional VAE to estimate the risk and migrate the anchor to reduce it .
Outcome: The proposed method improves the baseline classifier performance on PDTB v2.0 .
Fine-Grained Arabic Dialect Identification (C18-1)

Copied to clipboard

Challenge: Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification).
Approach: They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic.
Outcome: The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words.
Entity Enhancement for Implicit Discourse Relation Classification in the Biomedical Domain (2021.acl-short)

Copied to clipboard

Challenge: Discourse relation classification is a challenging task when the text domain is different from the standard Penn Discourse Treebank (PDTB) training corpus domain.
Approach: They propose to use the Biomedical Discourse Relation Bank to improve discourse relational argument representation by linking explicit instances of similar relations with a voting pipeline.
Outcome: The proposed model outperforms the pre-trained BioBERT model by 2% points.
Polarized Opinion Detection Improves the Detection of Toxic Language (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods for estimating polarized annotations are un-normalized and difficult to exploit in machine learning.
Approach: They propose a method for K-class text classification that exploits polarized texts in the dataset.
Outcome: The proposed method exploits polarized texts in a dataset and can improve classification performance.
Dynamically Updating Event Representations for Temporal Relation Classification with Multi-category Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models with independent classifiers for each TLINK category hinder from using the whole data.
Approach: They propose a temporal relation classification model that manages dynamic event representations across multiple TLINKs using multi-task learning to leverage the full size of data.
Outcome: The proposed model outperforms state-of-the-art models and two strong transfer learning baselines on English and Japanese data.
Cross-Target Stance Classification with Self-Attention Networks (P18-2)

Copied to clipboard

Challenge: In stance classification, the target on which the stance is made defines the boundary of the task, and a classifier is usually trained for prediction on the same target.
Approach: They propose a neural model that can generalize classifiers between different targets by finding useful information shared between relevant targets.
Outcome: The proposed model can generalize between relevant targets and find useful information shared between relevant target domains which improves generalization in certain scenarios.
Open Vocabulary Extreme Classification Using Generative Models (2022.findings-acl)

Copied to clipboard

Challenge: Extreme multi-label classification (XMC) aims at tagging content with subset of labels from an extremely large label set.
Approach: They propose a model that predicts a set of labels outside of the known vocabulary by using a loss-dependent loss-based loss-free model.
Outcome: The proposed model can predict labels outside the known vocabulary while performing on par with state-of-the-art solutions for known labels.
Stop Measuring Calibration When Humans Disagree (2022.emnlp-main)

Copied to clipboard

Challenge: Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., predictive probabilities are a good indication of how likely a prediction is to be correct.
Approach: They propose to measure calibration to human majority given inherent disagreements on tasks where humans inherently disagree about which class applies.
Outcome: The proposed measures capture key statistical properties of human judgements including class frequency, ranking and entropy.
Improving Formality Style Transfer with Context-Aware Rule Injection (2021.acl-long)

Copied to clipboard

Challenge: Existing language models pre-trained on large-scale corpora often do not work well for user-generated data where the language styles differ significantly from the mainstream text.
Approach: They propose a method for formality style transfer by injecting multiple rules into an end-to-end BERT-based encoder and decoder model.
Outcome: The proposed method outperforms existing rule-based FST approaches on tweet sentiment analysis tasks.
DiscSense: Automated Semantic Analysis of Discourse Markers (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for predicting discourse markers have been used to study link between markers and semantic relations .
Approach: They use a model trained to predict discourse markers between sentence pairs to predict plausible markers between sentences with a known semantic relation.
Outcome: The proposed method predicts markers between sentence pairs with a known semantic relation . the resulting dataset, named DiscSense, is publicly available .
Contextual Argument Component Classification for Class Discussions (2020.coling-main)

Copied to clipboard

Challenge: Argument mining systems often consider contextual information when training to perform tasks such as argument component identification, classification, and relation extraction.
Approach: They propose to incorporate speaker context and local discourse context into a model for classifying argument components in multi-party classroom discussions.
Outcome: The proposed model improves when varying context size and position . the results support the claim that context size is important .
Using a Penalty-based Loss Re-estimation Method to Improve Implicit Discourse Relation Classification (2020.coling-main)

Copied to clipboard

Challenge: inessential words are unintentionally misjudged as attention-worthy words and assigned heavier attention weights than should be.
Approach: They propose a penalty-based method to regulate the attention learning process by integrating penalty coefficients into the computation of loss by means of overstability of attention weight distributions.
Outcome: The proposed method improves on the Penn Discourse TreeBank corpus and is competitive compared to the state-of-the-art methods.
NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable.
Approach: They propose to penalize quick, intuitive "System 1" thinking by combining linguistic isolation with resistance to intuitive shortcuts to assess model's reasoning abilities.
Outcome: The proposed model penalizes quick, intuitive “System 1” thinking, isolating fundamental reasoning skills.
From legal to technical concept: Towards an automated classification of German political Twitter postings as criminal offenses (N19-1)

Copied to clipboard

Challenge: 'Network Enforcement Act' provides for a regulatory framework for 'illegal content' on social network platforms like Twitter or Facebook.
Approach: They propose a data annotation schema to determine whether a particular tweet could constitute a criminal offense and a binary classification schema to help with this.
Outcome: The proposed schema shows that the majority of offensive posts do not constitute a criminal offense and still contribute to public discourse.
Article Classification with Graph Neural Networks and Multigraphs (2024.lrec-main)

Copied to clipboard

Challenge: Existing and newly published articles require complex and complex pipelines to classify them into context-specific label taxonomies.
Approach: They propose to enrich Graph Neural Network pipelines with multi-graph representations that encode multiple signals of article relatedness as distinct edge types.
Outcome: The proposed methods improve the performance of a variety of GNN models compared to default graphs.
Animacy Denoting German Nouns: Annotation and Classification (2022.lrec-1)

Copied to clipboard

Challenge: Animacy detection is meant to distinguish words which denote humans from words used to denote non-humans.
Approach: They propose a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities.
Outcome: The proposed gold standard comprises almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities.
BERTweet’s TACO Fiesta: Contrasting Flavors On The Path Of Inference And Information-Driven Argument Mining On Twitter (2024.findings-naacl)

Copied to clipboard

Challenge: Argument mining is a challenging analytical task in the rich context of Twitter (now X).
Approach: They propose to optimize the embeddings of the BERTweet transformer for argument mining on Twitter and broader generalization across topics.
Outcome: The proposed approach improves classification and generalization across topics using a siamese network and a dataset.
What Speakers really Mean when they Ask Questions: Classification of Intentions with a Supervised Approach (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on hidden intentions of speakers in questions during meals is based on written or oral data, which are less easy to interpret.
Approach: They propose a typology of hidden intentions in questions asked during meals . they implement an automatic classification model based on annotated data and selected linguistic features.
Outcome: The proposed model is based on annotated data and features and evaluates its performance.
What Causes the Failure of Explicit to Implicit Discourse Relation Recognition? (2024.naacl-long)

Copied to clipboard

Challenge: Prior work claimed that explicit classifiers perform poorly in implicit scenarios . a label shift occurs after connectives are removed, but no empirical evidence supports this claim .
Approach: They propose to prove that the discourse relations expressed by some explicit instances will change when connectives disappear.
Outcome: The proposed methods outperform strong baselines on PDTB 2.0, PDTT 3.0, and the GUM dataset.
Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness (2021.emnlp-main)

Copied to clipboard

Challenge: Text style can reveal sensitive attributes of the author (e.g. age and race) to the reader, which can lead to privacy violations and bias in both human and algorithmic decisions based on text.
Approach: They propose a framework that obfuscates stylistic features of human-generated text through style transfer by automatically re-writing the text itself.
Outcome: The proposed framework obfuscates stylistic features of human-generated text through style transfer, by automatically re-writing the text itself.
BullyBench: Youth & Experts-in-the-loop Framework for Intrinsic and Extrinsic Cyberbullying NLP Benchmarking (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing youth-focused CB datasets lack conversational realism and ethical youth involvement with little or no evaluation of their social plausibility.
Approach: They propose a youth-in-the-loop dataset “BullyBench” that incorporates a structured intrinsic quality evaluation with experts-in the-looop (social scientists, psychologists, and content moderators) they perform extrinsic baseline evaluation by benchmarking encoder- and decoder-only language models for multi-class CB role classification.
Outcome: The proposed dataset is evaluated by a team of social scientists, psychologists, and content moderators to assess its quality, relevance, and coherence.
The Subject Annotations of the Danish Parliament Corpus (2009-2017) - Evaluated with Automatic Multi-label Classification (2022.lrec-1)

Copied to clipboard

Challenge: The interest in analysing and automatically processing large amounts of political data has increased in the past decades.
Approach: They address the semi-automatic annotation of subjects in the Danish Parliament Corpus (2009-2017) v.2 and describe multi-label classification experiments to verify the consistency of the subject annotation.
Outcome: The proposed method improves on the baseline classifier, which is a majority classifier.
Scientific Statement Classification over arXiv.org (2020.lrec-1)

Copied to clipboard

Challenge: a dataset of 1.2 million documents converted from the original submissions is available for supervised learning.
Approach: They propose a new classification task for scientific statements and a large-scale dataset for supervised learning.
Outcome: The proposed task achieves a 0.91 F1 score and a lexeme serialization for mathematical formulas.
A Retrospective Analysis of the Fake News Challenge Stance-Detection Task (C18-1)

Copied to clipboard

Challenge: 2017 Fake News Challenge Stage 1 (FNC-1) shared task addressed a stance classification task as a crucial first step towards detecting fake news.
Approach: They propose a new evaluation metric favoring the majority class, which can be easily classified, and propose stacked LSTM models that perform on par with the best systems, but is superior in predicting minority classes.
Outcome: The proposed evaluation metric favors the majority class, which can be easily classified, and overestimates the true discriminative power of the methods.
Automatic Animacy Classification for Romanian Nouns (2024.lrec-main)

Copied to clipboard

Challenge: Animacy is a semantic property of nouns that describes the quality of the noun's referent of being alive, sentient or volitional.
Approach: They propose a type-based binary classifier of Romanian nouns into the classes human/non-human using pre-trained word embeddings and animacy information derived from Romanian WordNet.
Outcome: The proposed classifiers perform well on the Romanian language and in a naturalistic setting.
D2GCLF: Document-to-Graph Classifier for Legal Document Classification (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods learn latent representations for each document by considering the semantics and themes of the documents.
Approach: They propose a document-to-graph classifier which extracts facts as relations between key participants in a law case and represents a legal document with four relation graphs.
Outcome: The proposed method outperforms the state-of-the-art methods on a real-world legal document dataset.
“Laughing at you or with you”: The Role of Sarcasm in Shaping the Disagreement Space (2021.eacl-main)

Copied to clipboard

Challenge: Detecting arguments in online interactions is useful to understand how conflicts arise and get resolved.
Approach: They propose to use a corpus annotated with argumentative moves and sarcasm to model sarcastic relationships using deep learning architectures.
Outcome: The proposed setup improves the argumentative relation classification task using deep learning architectures.
Native Language Identification in Texts: A Survey (2024.naacl-long)

Copied to clipboard

Challenge: Native language identification is the task of automatically identifying an author’s native language (L1) based on their second language production.
Approach: They present a survey of native language identification applied to texts . authors describe several text representations and computational techniques used in the task .
Outcome: The proposed task has been widely studied for both text and speech, particularly for L2 English due to the availability of suitable corpora.
Multi-label Categorization of Accounts of Sexism using a Neural Framework (D19-1)

Copied to clipboard

Challenge: Sexism manifests in blatant as well as subtle ways, authors say . existing work on sexism classification has limitations in terms of categories used . authors: categorization of accounts of sexist behavior can aid in countering sextism .
Approach: They propose a neural solution that can combine sentence representations with distributional and linguistic word embeddings.
Outcome: a new method outperforms deep learning and traditional methods by an appreciable margin . the proposed method outpersforms several deep learning as well as traditional baselines by an approval margin compared to baselines .
Automatic Classification of Neutralization Techniques in the Narrative of Climate Change Scepticism (2021.naacl-main)

Copied to clipboard

Challenge: neutralisation is used to justify lack of action or promote an alternative view of climate change . action on climate change has become an increasingly partisan issue with strong opposition voices discrediting scientists and spreading scepticism and misinformation.
Approach: They propose to use neutralisation techniques to introduce the problem to the nlp community and to collect manual annotations of neutralised techniques in text relating to climate change.
Outcome: The proposed models are supervised and semi-supervised by a team of researchers from the nlp and the ccsc.
Your fairness may vary: Pretrained language model fairness in toxic text classification (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained, bidirectional language models have revolutionized natural language processing research . authors show that focusing on accuracy measures alone can lead to models with wide variation in fairness characteristics .
Approach: They propose to use two post-processing methods to improve model fairness without retraining . they use pretrained language models of varying sizes on two toxic text classification tasks .
Outcome: The proposed methods improve model fairness without retraining . the results show that the fairness variation is more than just accuracy .
Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks (2023.eacl-main)

Copied to clipboard

Challenge: Disagreement can reflect different aspects of linguistic annotation, from annotators’ subjectivity to sloppiness or lack of context to interpret a text.
Approach: They propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks and manually label part of a Twitter dataset for offensive language detection in english following this taxonomies.
Outcome: The proposed taxonomy of disagreements in linguistic datasets can be used to assess how accurate tweets belonging to different disagreement categories can be classified as offensive or not.
Lifelong Learning of Hate Speech Classification on Social Media (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on automated hate speech classification assumes that the dataset is fixed and the classes are pre-defined.
Approach: They propose to use Variational Representation Learning and a load-balancing self-organizing inductive neural network to learn hate speech classification on social media.
Outcome: The proposed model improves on the lifelong learning techniques on social media.
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)

Copied to clipboard

Challenge: Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP).
Approach: They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays.
Outcome: The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc.
Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods of implicit discourse relation recognition (IDRR) focus on three aspects: enhancing discourse units representation, enhancing semantic interaction, and joint learning with other tasks.
Approach: They propose a joint model to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously.
Outcome: The proposed model achieves the best performance against several state-of-the-art systems on Chinese and English datasets.
How Many and Which Training Points Would Need to be Removed to Flip this Prediction? (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods to find St using brute-force are intractable.
Approach: They propose a fast approximation method to find St based on influence functions . they propose to identify a minimum subset of training data that one would need to remove .
Outcome: The proposed method can find St based on influence functions for simple classification models.
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)

Copied to clipboard

Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.
My Heart Skipped a Beat! Recognizing Expressions of Embodied Emotion in Natural Language (2024.naacl-long)

Copied to clipboard

Challenge: a new task is needed to recognize physical manifestations of emotions in natural language . physical manifestation of emotions affects not only our mental state but also our physical state .
Approach: They propose a task to recognize expressions of embodied emotion in natural language . they use body part mentions with human annotations to extract emotional manner expressions .
Outcome: The proposed model can train without gold data and improve performance with gold data.
IruMozhi: Automatically classifying diglossia in Tamil (2024.findings-naacl)

Copied to clipboard

Challenge: Literary Tamil is highly diglossic, with two very different registers in everyday use . Spoken Tamil is under-studied in modern NLP systems compared to Literary Tamil written in the Tamil script .
Approach: They present a human-translated dataset of parallel text in Literary and Spoken Tamil.
Outcome: The proposed model trains classifiers on the task of identifying which Tamil variety a text belongs to.
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies suggest using large language models to make tabular classifications . however, LLMs have been shown to exhibit harmful social biases based on stereotypes and inequalities present in society.
Approach: They propose to use large language models to make tabular classifications . they show that LLMs inherit biases from their training data .
Outcome: The proposed models exhibit harmful biases that reflect stereotypes and inequalities in society.
Training a Broad-Coverage German Sentiment Classification Model for Dialog Systems (2020.lrec-1)

Copied to clipboard

Challenge: Existing sentiment data sets are not available for sentiment analysis.
Approach: They propose to combine a German sentiment corpus with existing resources to train a general-purpose German sentiment classification model.
Outcome: The proposed model trains a general-purpose German sentiment classification model . the data set contains 5.4 million labelled samples .
The Craft of Selective Prediction: Towards Reliable Case Outcome Classification - An Empirical Study on European Court of Human Rights Cases (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing COC tasks prioritize high task performance over model reliability . however, large models exhibit overconfidence and Monte Carlo dropout methods produce reliable confidence estimates .
Approach: They conduct an empirical investigation into how various design choices affect the reliability of COC models within the framework of selective prediction.
Outcome: The proposed model is able to predict the outcome of a legal case based on the text of the case facts and is compared with other models using a pre-training corpus.
Towards Intelligent Clinically-Informed Language Analyses of People with Bipolar Disorder and Schizophrenia (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on social media data have limited the extent to which they can produce meaningful or generalizable conclusions.
Approach: They propose to use transcribed conversations with people with bipolar disorder and schizophrenia to create a large dataset of transcriptions.
Outcome: The proposed dataset extracts 100+ temporal, sentiment, psycholinguistic, emotion, and lexical features and establishes classification validity.
Cross-Domain Classification of Moral Values (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to identify moral values in text can be challenging for transferring knowledge between domains.
Approach: They compare a deep learning model with a domain-specific value classifier to find out whether it can transfer knowledge to new domains.
Outcome: The proposed model can generalize and transfer knowledge to novel domains, but introduce catastrophic forgetting.
BigNLI: Native Language Identification with Big Bird Embeddings (2024.lrec-main)

Copied to clipboard

Challenge: Native Language Identification (NLI) is a task that relies on time-consuming linguistic feature engineering and current transformer models are limited by input size.
Approach: They propose to train a logistic regression classifier which only uses Big Bird embeddings to overcome this limitation.
Outcome: The proposed method outperforms linguistic feature engineering models on the Reddit-L2 dataset and shows consistent out-of-sample and out-off-domain performance.
Confidence-based Ensembling of Perspective-aware Models (2023.emnlp-main)

Copied to clipboard

Challenge: Human label variability has been a topic of research in the field of NLP recently . Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fairer evaluation.
Approach: They propose a highly perspectivist model that exploits disagreements in annotations to capture the subjectivity encoded in the annotation process.
Outcome: The proposed model is validated on irony and hate speech detection scenarios in in-domain and cross-domain settings.
Are All the Datasets in Benchmark Necessary? A Pilot Study of Dataset Evaluation for Text Classification (2022.naacl-main)

Copied to clipboard

Challenge: Existing benchmark datasets contribute little to discriminating top-scoring systems, while those less used datasets exhibit impressive discriminative power.
Approach: They examine the distinguishability of benchmark datasets when comparing different systems . they find that existing benchmark dataset contribute little to discriminating top-scoring systems - whereas those less used datasets exhibit impressive discriminative power.
Outcome: The proposed datasets are released on DataLab.
Evaluating Evaluation Measures for Ordinal Classification and Ordinal Quantification (2021.acl-long)

Copied to clipboard

Challenge: Ordinal Classification (OC) tasks require ordinal classes, not nominal ones, to be evaluated.
Approach: They use data from the SemEval and NTCIR communities to clarify evaluation measures for Ordinal Classification and Ordinal Quantification tasks.
Outcome: The evaluation measures for Ordinal Classification (OC) and Ordinal Quantification (OQ) tasks are ordinal, not nominal.
Modularized Syntactic Neural Networks for Sentence Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for sentence classification use local information of sub-trees, but new models use global context .
Approach: They propose a tree-parallel mini-batch strategy for efficient training and predicting sentences . they propose to use syntax category labels to model sub-trees .
Outcome: The proposed model outperforms state-of-the-art tree-based methods on the sentence classification task.
Example-Driven Intent Prediction with Observers (2021.naacl-main)

Copied to clipboard

Challenge: Prior work has shown that BERT-like models attribute a significant amount of attention to the [CLS] token, which results in diluted representations.
Approach: They propose two approaches to improve generalizability of dialog system intent classification models by using observers and example-driven training.
Outcome: The proposed models achieve state-of-the-art on three intent prediction datasets in both the full data and few-shot settings.
Learning Features from Co-occurrences: A Theoretical Analysis (C18-1)

Copied to clipboard

Challenge: Existing theories for word classification and clustering are lacking.
Approach: They propose a theory that uses a function to represent a word by its co-occurrences with other words in context.
Outcome: The proposed model improves word classification and clustering by using multiple features.
Can Humans Identify Domains? (2024.lrec-main)

Copied to clipboard

Challenge: Textual domain is a crucial property within the Natural Language Processing community due to its effects on downstream model performance.
Approach: They examine the level of human disagreement and the relative difficulty of each annotation task by training classifiers to perform the same task.
Outcome: The authors show that human proficiency in identifying related intrinsic textual properties is low and that disagreements are high.
How Far Can It Go? On Intrinsic Gender Bias Mitigation for Text Classification (2023.eacl-main)

Copied to clipboard

Challenge: a growing interest in exploring how gender bias pertains in contextualized language models has been generated . intrinsic mitigation strategies and bias metrics have been proposed to mitigate gender bias in contextualised language models .
Approach: They propose to use different intrinsic bias mitigation strategies to mitigate gender bias in contextualized language models.
Outcome: The proposed probe shows that some mitigation techniques can hide gender bias . the probe also shows that not all mitigation techniques fool extrinsic bias despite their use .
kFolden: k-Fold Ensemble for Out-Of-Distribution Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies studying OOD detection in NLP often rely on external data to diversify model predictions.
Approach: They propose a framework which mimics OOD detection behavior without external data . they take text classification as an archetype and compare them to existing datasets .
Outcome: The proposed framework can resolve in- and out-distribution examples in a natural way using existing datasets.
Classifying Sluice Occurrences in Dialogue (L18-1)

Copied to clipboard

Challenge: Ellipsis is an important challenge for natural language processing systems, says a new paper . previous work on ellipsis focused on news data, but sluicing presents a challenge for dialogue systems .
Approach: They describe a corpus of 4100 sluice occurrences from the NYTimes Gigaword corpus . they build a classifier model to automatically classify slujce .
Outcome: The proposed corpus contains 4100 sluice occurrences, with an accuracy of 67% . the work will support empirical research into slujcing in dialogue systems .
Large-scale Exploration of Neural Relation Classification Architectures (D18-1)

Copied to clipboard

Challenge: Existing studies on relation classification have been limited to a very narrow range of datasets, making comparisons between approaches difficult.
Approach: They propose a multi-channel LSTM model combined with a CNN that takes advantage of all currently popular linguistic and architectural features.
Outcome: The proposed model achieves state-of-the-art on two datasets and provides direct insights into the challenges faced by language models on relation classification.
Using J-K-fold Cross Validation To Reduce Variance When Tuning NLP Models (C18-1)

Copied to clipboard

Challenge: a recent study shows that performance estimations are unstable and variable . this makes it difficult to use parameter tuning and model selection .
Approach: They propose to use a less variable CV method to evaluate performance . they propose lower choices of K than are typically seen in the NLP literature .
Outcome: The proposed method can be used for parameter tuning and performance estimation, but it is unstable and unstable.
Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Existing frameworks for imbalanced text classification can generate anchor instances for difficult samples . difficult samples are hard to classify as they are embedded into an overlapping semantic region with the majority class.
Approach: They propose a Mutual Information constrained Semantically Oversampling framework that generates anchor instances for difficult samples to help the backbone network determine the re-embedding position of a non-overlapping representation.
Outcome: The proposed framework can generate anchor instances to help classifiers achieve significant improvements over baselines on a variety of imbalanced text classification tasks.
Can We Identify Stance without Target Arguments? A Study for Rumour Stance Classification (2024.lrec-main)

Copied to clipboard

Challenge: Existing target-aware models underperform in cases where the context of the target is crucial.
Approach: They propose a framework to enhance reasoning with the targets and propose 'target-aware' models without awareness of the target.
Outcome: The proposed framework achieves state-of-the-art on two benchmark datasets.
CareCorpus: A Corpus of Real-World Solution-Focused Caregiver Strategies for Personalized Pediatric Rehabilitation Service Design (2024.lrec-main)

Copied to clipboard

Challenge: Pediatric rehabilitation services focus on functional skills and participation, defined as attendance and involvement in home, school, and community activities.
Approach: They propose to use a dataset of 780 real-world strategies written by caregivers to sort caregiver strategies for use in designing pediatric rehabilitation care plans.
Outcome: The proposed model can be used to sort caregiver strategies for use in designing pediatric rehabilitation care plans.
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification (2024.naacl-long)

Copied to clipboard

Challenge: Fallacy classification is a task of broad importance due to advances in deep learning and availability of more data.
Approach: They propose a new annotation scheme tailored for subjective NLP tasks and a method designed to handle subjectivity.
Outcome: The proposed approach integrates existing fallacy classification datasets with new ones.
Automatic Identification and Classification of Bragging in Social Media (2022.acl-long)

Copied to clipboard

Challenge: Bragging is a speech act employed to build a favorable self-image through positive statements about oneself.
Approach: They propose to use tweets annotated for bragging to build a model that can predict bragging with macro F1 up to 72.42 and 35.95 for binary and multi-class bragging classification tasks respectively.
Outcome: The proposed models predict bragging with macro F1 up to 72.42 and 35.95 in binary and multi-class classification tasks respectively.
PITA: Prompting Task Interaction for Argumentation Mining (2024.acl-long)

Copied to clipboard

Challenge: Argumentation mining (AM) aims to detect arguments and their inherent relations from textual compositions.
Approach: They propose a method to model the inter-relationships among three subtasks within a generative framework.
Outcome: The proposed method achieves state-of-the-art performance on two AM benchmarks.
Within-Between Lexical Relation Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for recognizing lexical-semantic relations between words are path-based and distributional.
Approach: They propose a novel Within-Between Relation model for recognizing lexical-semantic relations between words.
Outcome: The proposed model outperforms baselines across various benchmarks and is competitive and competitive.
C-Journal: A Journaling Application for Detecting and Classifying Cognitive Distortions Using Deep-Learning Based on a Crowd-sourced Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Cognitive distortions are negatively biased thinking patterns and erroneous self-statements resulting from and leading to logical errors in one’s internal reasoning.
Approach: They propose to provide an open-source labeled dataset of 14 cognitive distortions consisting of 34370 entries collected via crowd-sourcing, user questionnaires, and re-purposing emotions dataset from social media.
Outcome: The proposed model was built into a free journaling and mood-tracking mobile application that pinpoints cognitive distortions to the users.
Exploring the Limitations of Detecting Machine-Generated Text (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in the quality of the generation of text by large language models have spurred research into identifying machine-generated text.
Approach: They audit classification performance for detecting machine-generated text by evaluating on texts with varying writing styles.
Outcome: The proposed methods are highly sensitive to stylistic changes and complexity, and in some cases degrade entirely to random classifiers.
Emotion Classification by Jointly Learning to Lexiconize and Classify (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to identify emotions in short text are limited and lack coverage and inaccuracies when applied to informal short text.
Approach: They propose a novel emotional network to jointly learn sentence emotions and construct emotion lexicons which are dynamically adapted to a given context.
Outcome: The proposed model outperforms several approaches proposed in previous studies and achieves new state-of-the-art on the benchmark Twitter dataset.
Punctuation as Native Language Interference (C18-1)

Copied to clipboard

Challenge: Numerous aspects of written language have been studied for native language identification (NLI) but its impact has not been studied.
Approach: They propose to use punctuation marks as indicators of native language . they propose to apply them to native language identification .
Outcome: The proposed methods support the hypothesis that punctuation marks are persistent and robust indicators of the native language of the author, even when a high proficiency level in a non-native language is achieved.
”I Never Said That”: A dataset, taxonomy and baselines on response clarity classification (2024.findings-emnlp)

Copied to clipboard

Challenge: Equivocation and ambiguity in public speech are well-studied discourse phenomena . a new taxonomy aims to detect and classify response clarity in political interviews .
Approach: They propose a taxonomy that uses Large Language Models and human annotations to detect and classify response clarity in political interviews.
Outcome: The proposed taxonomy combines ChatGPT and human annotations to identify clarity in political questions . it provides a fine-grained taxonomies for evasion techniques related to unclear, ambiguous responses .
Exploring Discourse Structures for Argument Impact Classification (2021.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that discourse structures influence the persuasiveness of arguments.
Approach: They propose to fuse sentence-level structural discourse information with contextualized features derived from large-scale language models to investigate how discourse relations influence argument impact.
Outcome: The proposed model improves its backbone RoBERTa around 1.67%, compared with other models, but side effects are brought by other models.
Attention for Implicit Discourse Relation Recognition (L18-1)

Copied to clipboard

Challenge: Existing approaches to implicit discourse relation recognition reach F1 scores of 9.95% to 37.67% . a neural network exploits the strong correlation between pairs of words that implicitly signal a discourse relation.
Approach: They propose a neural network which exploits strong correlation between pairs of words . they use an encoder-decoder model with attention to detect a latent discourse relation .
Outcome: The proposed model outperforms state-of-the-art models on fine-grained classification and fine-granular classification while computing parameters without pooling and fully connected layers.
A Corpus of German Citizen Contributions in Mobility Planning: Supporting Evaluation Through Multidimensional Classification (2022.lrec-1)

Copied to clipboard

Challenge: Political authorities in democratic countries consult the public in order to allow citizens to voice their ideas and concerns on specific issues.
Approach: They propose a publicly-available corpus that includes citizen contributions from six mobility-related planning processes in five german municipalities.
Outcome: The proposed corpus includes several thousand citizen contributions from six mobility-related planning processes in five German municipalities.
Modeling Diagnostic Label Correlation for Automatic ICD Coding (2021.naacl-main)

Copied to clipboard

Challenge: Existing work built a binary prediction for each label independently, ignoring the dependencies between labels.
Approach: They propose a framework to capture the label correlation and train a reranking estimator to rescore the probability of each label set candidate generated by a base predictor.
Outcome: The proposed framework improves on the best-performing predictors on MIMIC datasets.
Implications of Using Internet Sting Corpora to Approximate Underage Victims (2021.findings-acl)

Copied to clipboard

Challenge: Existing systems for classification and triaging child exploitation cases require a high quality of data.
Approach: They propose to annotate a corpus of underage victim chats with convicted predators to compare their goals and tactics.
Outcome: The proposed model is based on a corpus of victim, vigilante, and LEO conversations with convicted predators.
Exploring Ordinality in Text Classification: A Comparative Study of Explicit and Implicit Techniques (2024.findings-acl)

Copied to clipboard

Challenge: Ordinal classification (OC) is a key task in natural language processing with applications in various domains such as sentiment analysis, rating prediction, and more.
Approach: They propose to tackle ordinal classification (OC) through the implicit semantics of the labels . they propose to use a classical explicit approach and an implicit approach that organically engages the semantics.
Outcome: The proposed methods are based on pre-trained language models and offer strategic recommendations based upon specific settings.
Integrating Tree Structures and Graph Structures with Neural Networks to Classify Discussion Discourse Acts (C18-1)

Copied to clipboard

Challenge: Existing models that analyze textual contents and discussion structures require understanding of textual content and discussion structure.
Approach: They propose a model that integrates discussion structures with neural networks to classify discourse acts.
Outcome: The proposed model improves accuracy and FB1 score by 1.5% compared to the previous best model.
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
Exploiting Noisy Data in Distant Supervision Relation Classification (N19-1)

Copied to clipboard

Challenge: Existing approaches to relation classification are noisy and time-consuming . RCEND uses noisy data to split noisy data into correctly and incorrectly labeled data .
Approach: They propose a framework to enhance relation classification by exploiting noisy data . they use an instance discriminator with reinforcement learning to split noisy data into correctly and incorrectly labeled data based on the noisy data.
Outcome: The proposed method outperforms the state-of-the-art models on relation classification . the proposed method is based on a semi-supervised learning method .
Which side are you on? Insider-Outsider classification in conspiracy-theoretic social media (2022.acl-long)

Copied to clipboard

Challenge: Social media is a breeding ground for threat narratives and related conspiracy theories.
Approach: They propose a new insider-outsider classification task that leverages pretrained language modeling to classify insiders and outsiders.
Outcome: The proposed model outperforms existing models and existing datasets in the classifying task.
Learning Disentangled Latent Topics for Twitter Rumour Veracity Classification (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches to rumour veracity classification relied on feature engineering.
Approach: They propose a model which disentangles the informational content of a tweet from the manner in which it is written.
Outcome: The proposed model disentangles the informational content of a tweet from the manner in which the information is written.
Event-Related Bias Removal for Real-time Disaster Events (2020.findings-emnlp)

Copied to clipboard

Challenge: Social media has become an important tool to share information about crisis events such as natural disasters and mass attacks.
Approach: They propose to train an adversarial neural model to remove latent event-specific biases and improve the performance on tweet importance classification.
Outcome: The proposed model removes event-specific biases and improves on tweet importance classification.
Is Probing All You Need? Indicator Tasks as an Alternative to Probing Embedding Spaces (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing probing tasks are designed to evaluate the information existing in representations by training a simple classification model.
Approach: They propose to use indicators to query embedding spaces for the existence of certain properties to determine whether a property exists in an embeddable space.
Outcome: The proposed indicators provide a more accurate picture of the information captured and removed compared to probes.
Leveraging Training Dynamics and Self-Training for Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Semi-supervised learning (SSL) is a promising technique for improving deep learning models when training data is scarce.
Approach: They propose a semi-supervised learning approach that leverages training dynamics of unlabeled data.
Outcome: The proposed method achieves an average increase in F1 score of 3.5% over baselines in low resource settings.
Scaling up Discourse Quality Annotation for Political Science (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotations on deliberative quality are time-consuming and suffer from class imbalance . ephd thesis: deliberation is not only the output of the decision making, but also the discussion that leads up to it.
Approach: They propose to use data augmentation techniques to improve deliberative quality predictions in a standard dataset.
Outcome: The proposed methods outperform classifiers based on linguistic features and argument quality annotations with or without data augmentation.
Contrast Sets for Stativity of English Verbs in Context (2022.coling-1)

Copied to clipboard

Challenge: Current methods for classifying verbs in context as dynamic or stative are limited to particular data sets.
Approach: They apply contrast set methodology to classify verbs in context as dynamic or stative . they create nearly 300 contrastive pairs by perturbing test set instances just enough to change their labels .
Outcome: The contrast set method is used to evaluate the performance of a model on a classifying task . the model performs worse on transformed examples than on human examples .
Annotating the French Wiktionary with supersenses for large scale lexical analysis: a use case to assess form-meaning relationships within the nominal lexicon (2025.coling-main)

Copied to clipboard

Challenge: Conducting large-scale empirical studies in lexical semantics remains an elusive goal for many languages lacking comprehensive semantic resources.
Approach: They propose to use the Princeton WordNet to enrich the French Wiktionary with general semantic classes, known as supersenses, using a limited amount of manually annotated data.
Outcome: The proposed method can be extended to other languages provided an electronic lexicon and manually annotated senses are available.
Corpus-based Identification of Verbs Participating in Verb Alternations Using Classification and Manual Annotation (2020.coling-main)

Copied to clipboard

Challenge: Verb alternations allow verbs to appear in a set of syntactically different constructions whose associated semantic frames are systematically related.
Approach: They use ENCOW and VerbNet data to train classifiers to predict the instrument subject alternation and the causative-inchoative alternation . they use count-based and vector-based features as well as perplexity-based language model features to reflect each alternation’s felicity by simulating it.
Outcome: The proposed approach reduces the required annotation effort by only presenting annotators with the highest-scoring candidates from the previous classification.
Enhancing Hierarchical Text Classification through Knowledge Graph Integration (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to hierarchical text classification are limited by lack of domain knowledge, which leads to mistakes in a variety of situations.
Approach: They propose a Knowledge-enabled Hierarchical Text Classification model which integrates knowledge graphs into HTC to address the knowledge limitations of traditional methods.
Outcome: The proposed model integrates knowledge graphs into the hierarchical text classification process, addressing the knowledge limitations of traditional methods.
Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Fine-grained classification tasks involve distinguishing between classes with subtle differences between them.
Approach: They analyse fine-grained text classification tasks by embedding class relationships into a contrastive objective function to help differently weigh the positives and negatives.
Outcome: The proposed model outperforms previous contrastive methods on emotion classification and sentiment analysis.
The Effect of Round-Trip Translation on Fairness in Sentiment Analysis (2021.emnlp-main)

Copied to clipboard

Challenge: Sentiment analysis systems exhibit sensitivity to protected attributes, while round-trip translation has been shown to normalize text.
Approach: They propose to use round-trip translation to normalize text to reduce the fairness gap between groups in sentiment analysis.
Outcome: The proposed method reduces the fairness gap between groups by up to 47%.
Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)

Copied to clipboard

Challenge: Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America.
Approach: They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data.
Outcome: The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data.
MHE: Code-Mixed Corpora for Similar Language Identification (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi .
Approach: They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines .
Outcome: The proposed corpus provides a language id at two levels: word and sentence.
Decoding the Narratives: Analyzing Personal Drug Experiences Shared on Reddit (2024.findings-acl)

Copied to clipboard

Challenge: Our study aims to develop a multi-level, multi-label classification model to analyze online user-generated texts about substance use experiences.
Approach: They propose a taxonomy to assess the nature of posts, including intended connections (Inquisition or Disclosure), subjects (e.g., Recovery, Dependency), and specific objectives (eg. relapse, quality, safety).
Outcome: The proposed model outperforms other models on annotated data and shows that topics such as Safety, Combination of Substances, and Mental Health see more disclosure, while discussions about physiological Effects focus on harm reduction.
Dialogue Act Classification with Context-Aware Self-Attention (N19-1)

Copied to clipboard

Challenge: Recent work in Dialogue Act classification has treated the task as a sequence labeling problem using hierarchical deep neural networks.
Approach: They propose a hierarchical deep neural network to model different levels of utterance and dialogue act semantics and use contextual dependencies to improve performance.
Outcome: The proposed model improves on the Switchboard Dialogue Act Corpus while maintaining high accuracy.
Seq2Emo: A Sequence to Multi-Label Emotion Classification Model (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for multi-label emotion classification are based on binary relevance and classifier chain (CC)
Approach: They propose a sequence-to-emotion approach which implicitly models emotion correlations in a bi-directional decoder.
Outcome: The proposed approach outperforms state-of-the-art methods on a SemEval’18 and GoEmotions dataset.
OTSeq2Set: An Optimal Transport Enhanced Sequence-to-Set Model for Extreme Multi-label Text Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Extreme multi-label text classification (XMTC) is a task of finding the most relevant subset labels from an extremely large label set.
Approach: They propose an autoregressive sequence-to-set model for XMTC tasks named OTSeq2Set.
Outcome: The proposed model outperforms other competitive baselines on 4 benchmark datasets.
Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance Weighting (2020.acl-main)

Copied to clipboard

Challenge: Recent research has found that text classification datasets contain certain unintended biases, such as text containing demographic identity-terms that are more likely to be abusive.
Approach: They propose a model-agnostic debiasing framework that recovers the non-discrimination distribution using instance weighting, which does not require extra resources or annotations apart from a pre-defined set of demographic identity-terms.
Outcome: The proposed framework alleviates the unintended biases without hurting models’ generalization ability.
Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Social media users are using images and text to voice opinions and share ideas.
Approach: They propose to use user comments to extract hinting features from user comments and explore them via self-training.
Outcome: The proposed framework improves on four social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection.
Multitask Learning for Emotionally Analyzing Sexual Abuse Disclosures (2021.naacl-main)

Copied to clipboard

Challenge: Prior work on identifying narratives related to sexual abuse disclosures did not consider this as an independent task.
Approach: They propose to identify narratives related to sexual abuse disclosures as a joint modeling task that leverages their emotional attributes through multitask learning.
Outcome: The proposed model leverages emotional attributes of textual conversations to identify narratives related to sexual abuse disclosures in homogeneous and heterogeneously settings.
Hierarchical CVAE for Fine-Grained Hate Speech Classification (D18-1)

Copied to clipboard

Challenge: Existing work on automated hate speech detection focuses on binary classification or on differentiating among a small set of categories.
Approach: They propose a method to discriminate among 40 hate groups of 13 different hate group categories.
Outcome: The proposed method outperforms discriminative models on a fine-grained hate speech classification task.
Cross-Cultural Transfer Learning for Text Classification (D19-1)

Copied to clipboard

Challenge: a large dataset is required to achieve competitive performance in most natural language tasks. large datasets are expensive, time consuming, and error-prone.
Approach: They propose a transfer-learning framework that leverages bilingual corpora for natural language text classification using no task-specific data.
Outcome: The proposed framework can achieve good performance on formality classification and sarcasm detection tasks without any task-specific labeled data.
Measuring Fairness of Text Classifiers via Prediction Sensitivity (2022.acl-long)

Copied to clipboard

Challenge: Existing fairness metrics are not yet available to measure the fairness of language processing systems.
Approach: They propose a new metric which measures fairness in machine learning models based on the model’s prediction sensitivity to perturbations in input features.
Outcome: The proposed metric can be linked with a specific notion of group fairness and individual fairness, and correlates well with humans’ perception of fairness.
Elvis vs. M. Jackson: Who has More Albums? Classification and Identification of Elements in Comparative Questions (2022.lrec-1)

Copied to clipboard

Challenge: Comparative Question Answering (cQA) is the task of providing accurate answers to questions . most question answering systems focus on answering factoid questions, but they fail at answering comparative questions in an efficient argumentative manner.
Approach: They propose two new open-domain datasets for identifying and labeling comparative questions . they use a binary classification task and an unsupervised sequence labeling task .
Outcome: The proposed datasets reach close-to-human results on a binary classification task with a neural model using ALBERT embeddings.
Distributional Term Set Expansion (L18-1)

Copied to clipboard

Challenge: Iterative term set expansion methods for distributional semantic models are used to label terms belonging to a sought after term set.
Approach: They compare iterative term set expansion methods for distributional semantic models to the Simple Margin method, an active learning approach to classification using Support Vector Machines.
Outcome: The proposed methods outperform centrality and classification based methods for distributional semantic models over five different term sets.
NaturalAdversaries: Can Naturalistic Adversaries Be as Effective as Artificial Adversaries? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work has explored adversarial example generation for natural language understanding tasks, but these examples are often unrealistic and diverge from the real-world data distributions.
Approach: They propose a framework for adversarial example generation that is effective at fooling a given classifier and a generative model based on the key tokens from the first stage.
Outcome: The proposed framework generalizes across domains and offers insights for future research on improving robustness of neural text classification models.
HeLI-OTS, Off-the-shelf Language Identifier for Text (2022.lrec-1)

Copied to clipboard

Challenge: Existing off-the-shelf language identification tools favor widely used languages, but Heli-OTS can be used to identify a large group of languages.
Approach: They introduce an off-the-shelf text language identification tool using the HeLI method . they compare the He LI-OTS language identifier with fastText on two different data sets .
Outcome: The proposed language identification tool is compared with fastText on two different data sets.
Leveraging Taxonomy and LLMs for Improved Multimodal Hierarchical Classification (2025.coling-main)

Copied to clipboard

Challenge: Multi-level Hierarchical Classification (MLHC) is a critical tool in modern data analysis.
Approach: They propose a taxonomy-embedded transitional LLM-agnostic framework for multimodality classification that leverages large language models to enforce consistency across hierarchical levels.
Outcome: The proposed framework improves on the MEP-3M dataset with various hierarchical levels compared to conventional models.
Neural Legal Judgment Prediction in English (P19-1)

Copied to clipboard

Challenge: Recent work on legal judgment prediction has focused on Chinese, but only feature-based models have been considered in English.
Approach: They propose a hierarchical version of BERT which bypasses BERT’s length limitation.
Outcome: The proposed model outperforms existing models in binary violation classification, multi-label classification and case importance prediction.
Topics to Avoid: Demoting Latent Confounds in Text Classification (D19-1)

Copied to clipboard

Challenge: Despite impressive performance on many text classification tasks, deep neural networks tend to learn frequent superficial patterns that are specific to the training data and do not always generalize well.
Approach: They propose a method that represents latent topical confounds and a model which “unlearns” confounding features by predicting both the label of the input text and the confound.
Outcome: The proposed model generalizes better and learns features indicative of the writing style rather than the content.
Exploring Distantly-Labeled Rationales in Neural Network Models (2021.acl-long)

Copied to clipboard

Challenge: Existing methods focus on distantly-labeled rationales, ignoring the potential important non-rationale words and not distinguishing the importance of different rationale words.
Approach: They propose two novel auxiliary loss functions to make better use of distantly-labeled rationales, which encourage models to maintain their focus on important words beyond labeled rationals (PINs) and alleviate redundant training on non-helpful rationale (NoIRs).
Outcome: The proposed methods outperform existing methods on two representative classification tasks while maintaining the ability to spread focus to other unlabeled important words.
On The Performance of Time-Pooling Strategies for End-to-End Spoken Language Identification (2020.lrec-1)

Copied to clipboard

Challenge: Language identification (LID) from speech is commonly tackled using similar approaches to those employed for speaker verification/recognition.
Approach: They propose to combine local descriptors and global descriptores into a single global description that can be used for downstream classification tasks.
Outcome: The proposed methods outperform well-known benchmark systems and previously results based on attention only.
DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing (2024.lrec-main)

Copied to clipboard

Challenge: DISRPT is a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing.
Approach: They present a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing that includes 13 languages and 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks.
Outcome: The DISRPT dataset includes data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks.
Comparing Text Representations: A Theory-Driven Approach (2021.emnlp-main)

Copied to clipboard

Challenge: Recent advances in NLP have been made by learning representations that transform complex tasks into simple classification tasks.
Approach: They propose a method to evaluate the compatibility between representations and tasks by fitting text features to specific characteristics of text datasets.
Outcome: The proposed model provides a calibrated, quantitative measure of the difficulty of a classification-based NLP task.
Cross Domain Classification of Education Talk Turns (2025.coling-main)

Copied to clipboard

Challenge: Prior research has focused on the annotation of conversational talk-turns within the classroom, offering a statistical analysis of the various types of discourse prevalent in these environments.
Approach: They examine the generalizability and transferability of text classifiers trained to predict classroom discourse across educational domains by accompanying each talk turn with dialog-level context.
Outcome: The proposed models exhibit high generalizability when training and test datasets originate from the same or similar domains.
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a new text classification framework for large language models addresses the problem of boundary ambiguity and inherent biases in LLMs.
Approach: They propose a two-stage classification framework for large language models to mitigate bottlenecks . their approach uses pairwise comparisons to efficiently narrow down options .
Outcome: The proposed framework reduces the number of options and improves on four datasets.
Multilingual and Multi-Aspect Hate Speech Analysis (D19-1)

Copied to clipboard

Challenge: Current research on hate speech analysis is oriented towards monolingual and single classification tasks.
Approach: They propose to use a multilingual multi-aspect hate speech analysis dataset to test current methods . they evaluate the dataset in various classification settings and discuss how to leverage annotations .
Outcome: The proposed dataset can be used to improve hate speech detection and classification in general.
A New Dataset for Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in natural language processing have lowered the barriers for people outside the NLP community to tap into the tools and resources applied to a variety of domain-specific applications.
Approach: They propose to annotate court transcripts from genocide-related cases using transformer-based approaches and to establish benchmarks for the task of paragraph identification of violence-related witness statements.
Outcome: The first annotated corpus of genocide-related court transcripts is aimed at providing a first reference corpus for the community and to establish benchmark performances using state-of-the-art transformer-based approaches.
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature.
Approach: They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders .
Outcome: The proposed evaluation protocol improves the existing framework and provides strong baseline results.
Implicit Sense-labeled Connective Recognition as Text Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for identifying implicit discourse relations are limited by the number of possible categories and sense labels.
Approach: They propose a method for identifying the sense label of an implicit connective between adjacent text spans by using an encoder-decoder model.
Outcome: The proposed method outperforms the conventional classification-based method on a shallow discourse parsing dataset.
Courtroom-LLM: A Legal-Inspired Multi-LLM Framework for Resolving Ambiguous Text Classifications (2025.coling-main)

Copied to clipboard

Challenge: Using a multi-LLM structure inspired by legal courtroom processes, we demonstrate that it can improve decision-making accuracy in ambiguous text classification scenarios.
Approach: They propose a legal-inspired multi-LLM structure that simulates a courtroom setting within LLMs and assigns roles similar to those of prosecutors, defense attorneys, and judges.
Outcome: The proposed model outperforms both single-LLM classifiers and simpler multi-LLMS setups in ambiguous text classification tasks.
Multi-Label Classification for Implicit Discourse Relation Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Prior research in discourse relation recognition has treated these instances as separate examples during training, with a gold-standard prediction matching one of the labels considered correct at test time.
Approach: They propose to use multiple labels to annotate an example when multiple relations are believed to hold simultaneously.
Outcome: The proposed frameworks don't depress performance for single-label prediction.
Exploiting a lexical resource for discourse connective disambiguation in German (2020.coling-main)

Copied to clipboard

Challenge: a connective lexicon can be a valuable resource for languages with limited PDTB-style annotations . connectives are usually understood to be ambiguous in two different ways .
Approach: They propose to augment a purely-empirical approach to connective identification and sense classification in German . they find that a connective lexicon can be a valuable resource for those languages with a large PDTB-style-annotated coprus .
Outcome: The proposed approach improves on published results and achieves an F1 score for German sense classification.
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback (2024.acl-long)

Copied to clipboard

Challenge: a growing body of work on learning from human feedback to align various aspects of machine learning systems with human values and preferences is focusing on the setting of fairness in content moderation.
Approach: They propose to use human feedback to determine how two comments should be treated in content moderation to learn about human values and preferences.
Outcome: The proposed approach is promising, as human preferences can often not be A: Some ladies like smaller men. B: Some men like smaller guys. Figure 1 shows that the proposed approach performs better for demographic intersections than a single classifier that gives equal weight to each annotation.
Native-like Expression Identification by Contrasting Native and Proficient Second Language Speakers (2020.coling-main)

Copied to clipboard

Challenge: a novel task of native-like expression identification is proposed by contrasting texts written by native speakers and those by proficient second language speakers.
Approach: They propose a task of native-like expression identification by contrasting texts written by native speakers and those by proficient second language speakers.
Outcome: The proposed method uncovers linguistically interesting usages distinctive of native speech.
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability.
Approach: They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria.
Outcome: The proposed system is based on 11 common aspects with different evaluation criteria.
A Deep Reinforced Sequence-to-Set Model for Multi-Label Classification (P19-1)

Copied to clipboard

Challenge: Multi-label classification (MLC) aims to assign multiple labels to each sample.
Approach: They propose a sequence-to-set model that is trained via reinforcement learning and rewards feedback independent of the label order.
Outcome: The proposed model outperforms baseline models and reduces sensitivity to label order.
Detecting Erroneously Recognized Handwritten Byzantine Text (2023.findings-emnlp)

Copied to clipboard

Challenge: Handwritten text recognition (HTR) produces textual output that contains errors, which are much higher than recognised printed text.
Approach: They investigate the properties of handwritten texts that lead post-correction systems to this adversarial behaviour in Byzantine Greek.
Outcome: The proposed model achieves an average precision score of 95% in Byzantine Greek and 97% in modern and ancient Greek.
A Context-based Framework for Modeling the Role and Function of On-line Resource Citations in Scientific Literature (D19-1)

Copied to clipboard

Challenge: Existing academic search engines cannot detect relevant papers where a resource is mentioned.
Approach: They propose a framework to model the role and function of on-line resource citations . they construct a dataset SciRes, which includes 3,088 manually annotated resource contexts based on a multi-task framework .
Outcome: The proposed model achieves the best results on both the classification task and recommendation task.
Classifying Referential and Non-referential It Using Gaze (D18-1)

Copied to clipboard

Challenge: a particular problem for anaphora resolution systems is the pronoun it, which can be used both referentially and non-referentially.
Approach: They use eye-tracking data to learn how humans perform disambiguation and use it to improve automatic classification.
Outcome: The proposed system outperforms a baseline and outperformed linguistic-based approaches.
Fact vs. Opinion: the Role of Argumentation Features in News Classification (2020.coling-main)

Copied to clipboard

Challenge: A 2018 study led by the Media Insight Project showed that most journalists think that their news organizations should clearly mark what is news reporting and what is commentary or opinion in order to combat fake news and gain public trust.
Approach: They propose to classify news articles into newsstories and opinion pieces using models that aim to sup-plement the article content representation with argumentation features.
Outcome: The proposed model outperforms linguistic features and improves on fine-tuned transformer-based models on data from publishers.
HSCNN: A Hybrid-Siamese Convolutional Neural Network for Extremely Imbalanced Multi-label Text Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to solve the data imbalance problem are limited in extremely imbalanced data.
Approach: They propose a hybrid approach which adapts general networks for head categories and few-shot techniques for tail categories.
Outcome: The proposed approach improves the performance of Single networks with diverse loss objectives on tail or entire categories.
Tribrid: Stance Classification with Neural Inconsistency Detection (2021.emnlp-main)

Copied to clipboard

Challenge: a new neural architecture can be used to classify stances on social media without relying on linguistic features.
Approach: They propose a neural architecture where the input also includes automatically generated negated perspectives over a given claim.
Outcome: The proposed model improves on the original input and removes doubtful predictions over the retained information.
Rethinking Attribute Representation and Injection for Sentiment Classification (D19-1)

Copied to clipboard

Challenge: Existing models that use text attributes to improve sentiment classification use text as a categorical feature.
Approach: They propose to represent attributes as chunk-wise importance weight matrices and consider four locations to inject attributes.
Outcome: The proposed method outperforms the state-of-the-art and outperformed previous models.
Predicting the Proficiency Level of Nonnative Hebrew Authors (2022.lrec-1)

Copied to clipboard

Challenge: a recent study shows that nonnative Hebrew learners can be accurately predicted from their essays . the proficiency level of nonnativ speakers is important for educational purposes .
Approach: They propose to use feature-based classifiers to accurately predict the proficiency level of nonnative Hebrew learners.
Outcome: The proposed classifiers can predict the proficiency level of nonnative Hebrew learners . the results are compared with human graders on a corpus of Hebrew essays .
Measuring Geographic Performance Disparities of Offensive Language Classifiers (2022.coling-1)

Copied to clipboard

Challenge: Recent work shows that text classifiers are biased regarding different languages and dialects.
Approach: They propose to use a dataset to examine whether language, dialect, and topical content vary across geographical regions to address these gaps.
Outcome: The proposed dataset includes 14 thousand examples across 15 cities and shows that current models do not generalize across locations.
Offensive Content Detection via Synthetic Code-Switched Text (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to detect offensive content in social media platforms are limited by the availability of labeled code-switched data.
Approach: They propose a method for generating synthetic code-switched offensive content data using human-generated data and a keyword classification baseline.
Outcome: The proposed algorithm can be used to generate synthetic code-switched offensive content data and train it on human-generated data.
Shami: A Corpus of Levantine Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world .
Approach: They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools.
Outcome: The proposed corpus is larger than existing corpora in terms of size, words and vocabularies.
Collocation Classification with Unsupervised Relation Vectors (P19-1)

Copied to clipboard

Challenge: Existing methods for relation classification are based on word embeddings, but they pose a greater challenge than syntactic and semantic relations.
Approach: They propose a distributional landscape based on word embeddings as a suitable basis for relation classification of collocations . they also conduct experiments on a subset of this benchmark .
Outcome: The proposed dataset is compared to the well known DiffVec dataset and shows that it is more efficient than the standard methods.
Detecting Minority Arguments for Mutual Understanding: A Moderation Tool for the Online Climate Change Debate (2022.coling-1)

Copied to clipboard

Challenge: polarized topics such as climate change present challenges for moderators and researchers.
Approach: They propose a moderation tool to support moderators in promoting mutual understanding in the online climate change debate by training classifiers to label incoming posts for the arguments they entail and using active learning to supplement the training data with rare arguments.
Outcome: The proposed method can be part of the toolkit for moderators struggling with polarized topics such as climate change.
The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Gender-fair language fosters inclusion by addressing all genders or using neutral forms.
Approach: They present a dataset that provides high-quality reformulations for German text classification . they find substantial label flips, reduced prediction certainty, and altered attention patterns .
Outcome: The proposed dataset provides high-quality reformulations for German text classification . it finds label flips, reduced prediction certainty, and significantly altered attention patterns .
Exploring BERT-Based Classification Models for Detecting Phobia Subtypes: A Novel Tweet Dataset and Comparative Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Phobias are characterized by an intense and irrational fear of specific objects, situations, or activities despite there being no real risk or only a minor threat involved.
Approach: They propose to use a dataset of 811,569 English tweets from user timelines spanning 102 phobia subtypes over six months to classify users into 65 specific phobias.
Outcome: The proposed dataset includes 47,614 self-diagnosed phobia users and a high f1 score for binary classification and multi-class classification.
Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed in decision-making tasks where accuracy and reliable confidence estimates are essential.
Approach: They propose a calibration-aware reinforcement learning formulation that directly adjusts decision-token probabilities.
Outcome: The proposed model preserves RLVR’s accuracy level while mitigating overconfidence, reducing ECE scores up to 9 points.
Impromptu Cybercrime Euphemism Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for detecting euphemisms are ineffective in impromptu euphorism detection . Existing approaches for e-mail detection are limited to word-level ephemismals .
Approach: They propose a framework for impromptu euphemism detection that integrates context augmentation and multi-round iterative training to better predict the actual meaning of a masked token.
Outcome: The proposed framework improves 76-fold over the previous state-of-the-art euphemism detector.
Classifying Dyads for Militarized Conflict Analysis (2021.emnlp-main)

Copied to clipboard

Challenge: Existing research examines the origins of militarized conflict by examining bi-lateral relationships between entity pairs and multi-lateral relations among multiple entities.
Approach: They propose to use Wikipedia to model dyadic and systemic causes to compare their correlations with conflict between two entities.
Outcome: The proposed graphs show that Wikipedia articles of allies are semantically more similar than enemies.
Figurative Language Processing: A Linguistically Informed Feature Analysis of the Behavior of Language Models and Humans (2023.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed a growing interest in investigating what Transformer-based language models (TLMs) actually learn from training data.
Approach: They propose to use a black-box TLM and two intrinsically transparent white-box models to investigate the performance of figurative language models on sarcasm, similes, idioms, and metaphors.
Outcome: The proposed models perform better than other models on figurative language classification tasks.
Offensive Language Identification in Greek (2020.lrec-1)

Copied to clipboard

Challenge: a gap in the literature on offensive language has been addressed with studies on Spanish, Hindi, and German.
Approach: They present a Greek annotated dataset for offensive language identification . it contains 4,779 tweets annotating offensive and not offensive posts from Twitter . they evaluate several computational models trained and tested on the dataset .
Outcome: The proposed dataset contains 4,779 tweets annotated as offensive and not offensive . the authors show that the proposed dataset is similar to the OLID dataset for English .
Estimating Confidence of Predictions of Individual Classifiers and TheirEnsembles for the Genre Classification Task (2022.lrec-1)

Copied to clipboard

Challenge: Genre identification is a kind of non-topic text classification. genre is defined as a functional space.
Approach: They propose to use SOTA to identify genres in non-topic texts . genres are functional and cannot be expressed just by some keywords .
Outcome: The proposed models show that they perform better than their individual models in large datasets.
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection (2020.acl-main)

Copied to clipboard

Challenge: Word embeddings, pre-trained language models, and deep learning methods are becoming effective for text classification.
Approach: They propose a method for removing information from neural representations using null-space projection.
Outcome: The proposed method mitigates bias in word embeddings and increases fairness in multi-class classification.
2kenize: Tying Subword Sequences for Chinese Script Conversion (2020.acl-main)

Copied to clipboard

Challenge: Traditional Chinese character conversion is a common step in Chinese NLP but current methods do not take into account that a simplified Chinese character can correspond to multiple traditional characters.
Approach: They propose a model that can disambiguate between mappings and convert between the two scripts by using subword segmentation and two language models.
Outcome: The proposed model outperforms previous Chinese Character conversion approaches by 6 points in accuracy.
Discriminatively-Tuned Generative Classifiers for Robust Natural Language Inference (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work has shown advantages of generative classifiers in terms of data efficiency and robustness.
Approach: They propose a generative classifier for natural language inference (NLI) they compare it to discriminative models and large-scale pretrained models like BERT .
Outcome: The proposed classifier outperforms discriminative and pretrained baselines across several challenging NLI experimental settings, including small training sets, imbalanced label distributions, and label noise.
Comparing Methods for Measuring Dialect Similarity in Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: a coarse-grained transcription of speech is sufficient to replicate dialectal boundaries, but it can be generalised over by an automatic method.
Approach: They propose to use two different methods to measure dialect similarity in Norwegian . they use the Levenshtein method and the neural long short term memory algorithm . the paper shows that coarse-grained transcriptions of speech can generate dialect maps .
Outcome: The proposed method can generalise over coarse-grained transcriptions, but it needs a large dataset . the proposed method is compared with canonical maps found in the dialect literature .
‘Aye’ or ‘No’? Speech-level Sentiment Analysis of Hansard UK Parliamentary Debate Transcripts (L18-1)

Copied to clipboard

Challenge: Transcripts of UK parliamentary debates are difficult for human readers to process due to the large quantity of textual data and the specialised language used.
Approach: They propose to use annotated sentiment labels and labels derived from speakers' votes to classify the sentiment polarity of speakers as being either positive or negative towards motions proposed in the debates.
Outcome: The proposed model outperforms existing models on a dataset of parliamentary debate transcripts using textual and contextual features.
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training (2024.emnlp-main)

Copied to clipboard

Challenge: Neural networks are increasingly prevalent across a wide range of applications, driving significant advancements in fields such as natural language processing, computer vision, and beyond.
Approach: They propose an end-to-end differentiable training paradigm for stable training of a rationalized transformer classifier.
Outcome: The proposed model is capable of classifying a sample and scoring input tokens without any explicit supervision and produces class-wise rationales without instabilities.
From Text to Source: Results in Detecting Large Language Model-Generated Content (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate human-like text, but have ethical and misuse concerns.
Approach: They evaluate whether a classifier trained to distinguish between source and target LLMs can detect text from an LLM without further training.
Outcome: The proposed method detects text from target LLMs without further training.
Generating Hard-Negative Out-of-Scope Data with ChatGPT for Intent Classification (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have not studied the performance of intent classifiers against hard-negative out-of-scope utterances.
Approach: They propose to generate hard-negative OOS data using ChatGPT and evaluate them against three benchmark intent classifiers.
Outcome: The proposed method improves classifiers' robustness against hard-negative out-of-scope utterances and general OOS data.
MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems (2026.acl-long)

Copied to clipboard

Challenge: MONETA is the first multimodal industry classification benchmark with text and geospatial sources.
Approach: They propose a multimodal industry classification benchmark using text and geospatial sources.
Outcome: The proposed model increases the accuracy of the existing models by 22.80%.
GERMS-AT: A Sexism/Misogyny Dataset of Forum Comments from an Austrian Online Newspaper (2024.lrec-main)

Copied to clipboard

Challenge: sexism/misogyny dataset extracted from comments of online forum of newspaper . corpus of 8 000 comments annotated with 5 levels of sexist/mistoginist .
Approach: They present a sexism/misogyny dataset extracted from comments of an online forum of an Austrian newspaper.
Outcome: The results show that the corpus of comments is sexist/misogynistic and has 5 levels of sexism/mistoginess.
Argument-based Detection and Classification of Fallacies in Political Debates (2023.emnlp-main)

Copied to clipboard

Challenge: Fallacies are arguments that employ faulty reasoning, causing inaccurate conclusions and invalid inferences . ad hominem fallacy is one of the most common fallacy labels used in political debates despite its use in many scenarios .
Approach: They extend the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments by incorporating the most recent Trump-Biden debate.
Outcome: The proposed method extends the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments .
Czech Text Document Corpus v 2.0 (L18-1)

Copied to clipboard

Challenge: a corpus of text documents for automatic document classification in Czech is presented . paper aims to facilitate a straightforward comparison of document classification approaches on Czech data .
Approach: This paper introduces a collection of text documents for automatic document classification in Czech language.
Outcome: The proposed corpus is based on the Czech news agency's real newspaper articles . it is used for evaluation of multi-label document classification approaches .
The Linguistic Category Model in Polish (LCM-PL) (L18-1)

Copied to clipboard

Challenge: a new version of the Linguistic Category Model (LCM) dictionary for the Polish language is available for use and integrates with the Polish WordNet.
Approach: They propose to use a dictionary that is annotated manually in its most important parts . they propose to add more manually annotating senses and increase quality of automated annotations .
Outcome: The proposed dictionary is the first widely usable version of the resource . it will have more manually annotated senses and more automated annotations .
Returning the N to NLP: Towards Contextually Personalized Classification Models (2020.acl-main)

Copied to clipboard

Challenge: a recent study shows that NLP models treat language as universal, but that it is based on sociolinguistic research.
Approach: They propose to incorporate user-dependent, contextual personal and social aspects into neural NLP models by means of socially contextual personalization.
Outcome: The proposed approach could be adapted to better personalize the language of users . it outlines a possible direction to incorporate these aspects into neural NLP models .
Language Identification for Austronesian Languages (2022.lrec-1)

Copied to clipboard

Challenge: This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages.
Approach: They compare a classifier based on skip-gram embeddings with other methods . they then increase the number of non-Austronesian languages to 800 to evaluate their performance .
Outcome: The proposed model improves on the previous methods for low- and under-resourced languages in the Pacific region.
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails (2025.findings-acl)

Copied to clipboard

Challenge: Existing guardrails rely on rule-based filtering or single-pass classification, limiting their ability to handle nuanced safety violations.
Approach: They propose a critique-augmented guardrail model that distills knowledge from high-capacity LLMs by generating structured critiques alongside safety labels.
Outcome: The proposed model outperforms existing guardrail models on multiple safety benchmarks and achieves the highest average F1 and AUPRC.
Thirumurai: A Large Dataset of Tamil Shaivite Poems and Classification of Tamil Pann (2022.lrec-1)

Copied to clipboard

Challenge: Thirumurai is a collection of Tamil Shaivite poems dating back to the Hindu revival period . a large dataset containing all the Thirumuru poems is under-resourced .
Approach: They propose to use transformers to classify the Tamil Pann and author of each poem . they propose to train models on petabytes of data, such as the common crawl data .
Outcome: The proposed dataset contains all the Thirumurai poems and classifies the Pann and author of each poem using transformer based architectures.
Discourse Context Predictability Effects in Hindi Word Order (2022.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that information status, dependency length, and syntactic surprisal influence word order preferences, but the role of discourse predictability is underexplored in the literature.
Approach: They propose to use Hindi-Urdu Treebank corpus to build a classifier to predict which sentences actually occurred in the corpus against artificially generated distractors.
Outcome: The proposed classifier predicts which sentences occur in the Hindi-Urdu Treebank corpus against artificial distractors.
How Diplomats Dispute: The UN Security Council Conflict Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Until now, there has been little work on how to formalize conflicts in a diplomatic setting.
Approach: They present a corpus of 87 UNSC speeches that are annotated for conflicts and demonstrate the difficulty when dealing with diplomatic language.
Outcome: The proposed method demonstrates that diplomatic language is complex and often implicit along various dimensions.
Feature Projection for Improved Text Classification (2020.acl-main)

Copied to clipboard

Challenge: In sentiment classification, there are some good features that are indicative of class labels, but there are also many common features that do not discriminate for classification.
Approach: They propose to project existing features into the orthogonal space of the common features and make them more discriminative for classification.
Outcome: The proposed method improves CNN, RNN, Transformer, and Bert based text classification and obtains markedly better results.
What time is it? Temporal Analysis of Novels (2020.emnlp-main)

Copied to clipboard

Challenge: a novel based on the flow of time provides a framework for understanding the text . a computational approach to annotate a book's lines with wall clock times is needed to understand the flow through time.
Approach: They propose to annotate each line of a book with wall clock times . they use a data set of hourly time phrases from 52,183 fictional books .
Outcome: The proposed method improves upon baselines by over two hours and can partition a book into segments that correspond to a particular time-of-day.
Humanitarian Corpora for English, French and Spanish (2024.lrec-main)

Copied to clipboard

Challenge: et al., a leading database of humanitarian documents, compiled with ReliefWeb reports . documents selected with language identification and noise reduction techniques . authors present corpora of English, French and Spanish humanitarian documents .
Approach: They present three corpora of English, French and Spanish humanitarian documents compiled with ReliefWeb reports . documents were tokenized, lemmatized, tagged by part of speech, and enriched with metadata . authors propose a project to develop a humanitarian dictionary with a focus on conceptual variation .
Outcome: The corpora were compiled to satisfy the research needs of the Humanitarian Encyclopedia project with a focus on conceptual variation.
KOLD: Korean Offensive Language Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Recent directions for offensive language detection focus on English and do not transfer well to other languages because of cultural and linguistic differences.
Approach: They present a Korean offensive language dataset annotated with offensive language comments . they use the comments as training data for Korean BERT and RoBERTa models .
Outcome: The proposed model improves offensiveness detection, target classification, and span detection while having room for improvement for target group classification and span prediction.
An information theoretic view on selecting linguistic probes (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in NLP tasks require a question of how much linguistic knowledge is encoded in neural networks.
Approach: They propose to use diagnostic classifiers to perform supervised classification from internal representations.
Outcome: Empirically, the two proposed criteria lead to results that agree with each other.
Identifying Fine-grained Depression Signs in Social Media Posts (2024.lrec-main)

Copied to clipboard

Challenge: Currently, most studies focus on a binary classification setup or on pre-established resources.
Approach: They evaluated machine learning techniques to model 21 depression signs in social media posts from Brazilian undergraduate students.
Outcome: The proposed methods struggle to classify the majority of depression signs on social media posts, compared with the majority on the social media sites.
Tackling Irony Detection using Ensemble Classifiers (2022.lrec-1)

Copied to clipboard

Challenge: Automated approaches to irony detection still fall short of what one would consider desirable performance.
Approach: They propose to use transformer-based approaches to automate irony detection in social media . they propose to augmentation training data to address the binary and fine-grained problem .
Outcome: The proposed methods improve performance over baselines and are not decisive for good results.
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions.
Approach: They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language .
Outcome: The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language .
Deep One-Class Hate Speech Detection Model (2022.lrec-1)

Copied to clipboard

Challenge: Existing approaches to hate speech detection neglect distinct attributes of hate speeches from other sentimental types such as “aggressive” and “racist”.
Approach: They propose a one-class model where the detection classifier is trained with hate-class samples only.
Outcome: The proposed model outperforms existing models with four benchmark datasets and shows that it performs better than existing models.
Inferring Social Media Users’ Mental Health Status from Multimodal Information (2020.lrec-1)

Copied to clipboard

Challenge: In the United States alone, one in every four adults suffers from a mental health condition, making mental health a pressing concern.
Approach: They propose to use multimodal cues present in social media posts to predict mental health status by analyzing language, visual, and metadata cue data.
Outcome: The proposed approach improves the performance of the classification task compared to using one modality at a time and can provide important cues into a user’s mental status.
Identifying Physical Object Use in Sentences (2022.emnlp-main)

Copied to clipboard

Challenge: Prior research has focused on learning the prototypical functions of physical objects . but many sentences refer to objects even when they are not used .
Approach: They propose a task that determines whether a physical object mentioned in a sentence was used or likely will be used.
Outcome: The proposed model exploits data augmentation methods and FrameNet to fine-tune a pre-trainedmodel.
Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper (2023.findings-emnlp)

Copied to clipboard

Challenge: Several platforms currently rate news outlets according to their political bias, but they are not ready for use in the near future.
Approach: They use ratings of authentic news outlets to create a multilingual corpus of news with coarse stance annotations along with automatically extracted topic annotations.
Outcome: The proposed model can identify the editorial line of most unseen newspapers in English, German, Spanish and Catalan.
InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks (2024.lrec-main)

Copied to clipboard

Challenge: InfFeed uses influence functions to compute the influential instances for a target instance.
Approach: They propose an apparatus that uses influence functions to compute the influential instances for a target instance.
Outcome: The proposed model outperforms the state-of-the-art baselines by 4% for hate speech classification, 3.5% for stance classification, and 3% for irony and 2% for sarcasm detection.
On Classifying whether Two Texts are on the Same Side of an Argument (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to same side stance classification (S3C) require domain knowledge and semantic inference to solve the task.
Approach: They propose to use same side stance classification to predict whether two arguments argue for the same stance for a given pair of arguments.
Outcome: The proposed model fails to generalize both within and across topics and domains when adjusting the sampling strategy to a more adversarial scenario.
Deep learning-based end-to-end spoken language identification system for domain-mismatched scenario (2022.lrec-1)

Copied to clipboard

Challenge: Domain mismatch is a critical issue when it comes to spoken language identification.
Approach: They evaluated a set of cross-domain language identification trials using a dataset from the Oriental Language Recognition (OLR) Challenge 2021 .
Outcome: The proposed architectures and deep learning strategies have shown good performance in cross-domain speaker verification tasks.
Token Prediction as Implicit Classification to Identify LLM-Generated Text (2023.emnlp-main)

Copied to clipboard

Challenge: a novel approach for identifying large language models (LLMs) involved in text generation is proposed . instead of adding an additional classification layer, we reframe the classification task as a next-token prediction task .
Approach: They propose a novel approach for identifying large language models involved in text generation . instead of adding an additional classification layer, they reframe the task as a next-token prediction task .
Outcome: The proposed method performs exceptionally well in the text classification task . it can distinguish distinctive writing styles among various LLMs even without an explicit classifier.
BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for multimodal learning are difficult to scale beyond two modalities and lack resolution for instance-level control.
Approach: They propose a bi-level weighting framework that combines instance-level Kullback-Leibler divergence and modality-level mutual information to dynamically adjust modality importance during training.
Outcome: The proposed method significantly improves regression performance and multiclass classification accuracy.
Learning under Label Proportions for Text Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Using a set of data samples with ground truth labels, learning from label proportions is infeasible in many cases.
Approach: They propose a setup where the data is provided in an aggregate form called bags and only the proportion of samples belonging to each class as the ground truth.
Outcome: The proposed model achieves better results than baselines in almost 87% of the configurations.
Theory of Mind in Freely-Told Children’s Narratives: A Classification Approach (2023.findings-acl)

Copied to clipboard

Challenge: Language and Theory of Mind (ToM) competences are often studied with younger children and standardized tests, but as both are social competences, data and methods with higher ecological validity are critical.
Approach: They leveraged a corpus of 442 freely-told stories by Dutch children aged 4-12 to study language and ToM with NLP-tools.
Outcome: The proposed method is robust relative to the complexity of the task for humans and is consistent with previous studies.
JL-Hate: An Annotated Dataset for Joint Learning of Hate Speech and Target Detection (2024.lrec-main)

Copied to clipboard

Challenge: Existing data resources for the detection of hate speech focus on text sequence classification, but the target of hateful content is lacking.
Approach: They propose a tweet dataset for the task of joint learning of hate speech detection and target detection called JL-Hate.
Outcome: The proposed dataset performs similar tasks to the existing datasets in sequence and token classification tasks.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech (2025.emnlp-main)

Copied to clipboard

Challenge: a new computational model for political delegitimization discourse is proposed for analysis of democratic discourse . we identify the importance of PDD as a powerful tool in political competition .
Approach: They propose a computational classification pipeline for political delegitimization discourse . they annotate a Hebrew-language corpus of 10,410 sentences from parliamentary speeches, facebook posts and leading news outlets .
Outcome: The proposed model achieves an F1 of 0.74 for binary detection and a macro-F1 of 0.6 for classification of delegitimization characteristics.
AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text (2025.findings-emnlp)

Copied to clipboard

Challenge: Domain adaptive pre-training and task-adaptive pre- training (TAPT) are popular methods to reduce this bias for low-resource languages, but they have not been explored for African multilingual encoders.
Approach: They propose a large-scale social media and news domain corpus for continual pre-training on African languages.
Outcome: The proposed methods improve performance on three subjective tasks, including sentiment analysis, multi-label emotion, and hate speech classification, while TAPT improves performance on other related tasks.
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes (2024.lrec-main)

Copied to clipboard

Challenge: Currently, sentiment analysis is a widely employed text classification task that involves extracting the sentiment expressed by individuals towards a variety of entities.
Approach: They propose to use KazSAnDRA to automate Kazakh sentiment analysis by developing and evaluating four machine learning models for polarity and score classification.
Outcome: The proposed dataset is the first and largest publicly available dataset of its kind.
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks (2024.findings-acl)

Copied to clipboard

Challenge: Memorisation in neural models is concerned due to overfitting and privacy concerns . a dominant hypothesis based on image classification is that lower layers learn generalisable features and deeper layers specialise and memorise.
Approach: They apply 4 techniques to localise and edit models' memories.
Outcome: The proposed method shows that memorisation is a gradual process rather than a localised one.
Annotation-Inspired Implicit Discourse Relation Classification with Auxiliary Discourse Connective Generation (2023.acl-long)

Copied to clipboard

Challenge: Discourse connectives are words or phrases that signal the presence of a discourse relation.
Approach: They propose a model that generates discourse connectives between arguments and predicts discourse relations based on the generated connectives.
Outcome: The proposed model outperforms baselines on three datasets and is highly accurate.
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Adapting cultural values in Large Language Models presents significant challenges due to biases and data limitations.
Approach: They propose to augment World Values Survey (WVS) data with encyclopedic and scenario-based cultural narratives from Wikipedia and NormAd to address these limitations.
Outcome: The proposed approach enhances cultural distinctiveness and improves classification performance across cultures.
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification (2024.emnlp-main)

Copied to clipboard

Challenge: a novel dataset of text-embedded images associated with the LGBTQ+ Pride movement is presented in this paper . a new framework for analyzing text-based images is proposed to address this challenge .
Approach: They propose a new dataset for machine learning that includes hate, targets of hate, stance, humor and a framework for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.
Outcome: The proposed framework achieves superior performance on two real-world datasets.
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on toxicity and polarization as separate problems . extreme polarizing deepens divisions, often leading to hostility and fragmentation .
Approach: They propose to use a multi-label Indonesian dataset annotated for toxicity, polarization, and annotator demographic information to study polarizing language and toxicity.
Outcome: The proposed dataset shows that polarization cues improve toxicity classification and vice versa.
MentalRiskES: A New Corpus for Early Detection of Mental Disorders in Spanish (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on the prevalence of mental disorders on the Web are limited to the English language.
Approach: They propose to use user messages posted on Telegram groups to annotate the corpus for natural language processing and to conduct experiments on text classification and regression.
Outcome: The proposed corpus contains over 1,300 subjects with more than 45,000 messages posted in different public Telegram groups.
Vulnerability of LLMs to Vertically Aligned Text Manipulations (2025.acl-long)

Copied to clipboard

Challenge: Recent research shows that vertical text input significantly degrades the accuracy of large language models (LLMs) in text classification tasks.
Approach: They investigate the impact of vertical text input on the performance of LLMs . they find that chain of thought reasoning does not help LLM recognize vertical input .
Outcome: The proposed model can significantly mislead models, posing a risk of bypassing detection in real-world scenarios involving harmful or sensitive information.
Fair Text Classification with Wasserstein Independence (2023.emnlp-main)

Copied to clipboard

Challenge: a novel method for mitigating group fairness in text classification remains an open challenge . a recent paper focuses on the problem of group fair in the domain of natural language processing .
Approach: They propose a method for mitigating group fairness in neural text classification . they use adversarial training to induce Wasserstein independence between representations .
Outcome: The proposed method exhibits a comparable or better fairness-accuracy trade-off compared to existing methods.
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web.
Approach: They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances.
Outcome: The proposed classifier augments training data with automatically-generated GPT-3 completions.
Debiasing Online Preference Learning via Preference Feature Preservation (2025.findings-acl)

Copied to clipboard

Challenge: Recent preference learning frameworks for large language models (LLMs) simplify human preferences with binary pairwise comparisons and scalar rewards.
Approach: They propose a preference feature preservation framework that preserves the distribution of human preference features and maps them throughout the online preference learning process.
Outcome: The proposed framework maintains the distribution of human preference features and utilizing such rich signals throughout the online preference learning process.
Act2P: LLM-Driven Online Dialogue Act Classification for Power Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on explicit utterance functions, overlooking the implicit power dynamics embedded in dialogue.
Approach: They propose an online Dialogue Act Classification and Dynamic Power Analysis framework based on large language models to integrate dialogue act classification with power quantification.
Outcome: The proposed framework outperforms existing methods in online scenarios and shows that dialogue power is distributed and dynamically transferred.
Argument Relation Classification through Discourse Markers and Adversarial Training (2024.emnlp-main)

Copied to clipboard

Challenge: Argument relation classification (ARC) identifies supportive, contrasting and neutral relations between argumentative units.
Approach: They propose an argument relation classifier that integrates knowledge of discourse markers into a pre-trained RoBERTa model.
Outcome: The proposed model outperforms existing methods and learns discriminative sentence embeddings supporting the task.
GAttention: Gated Attention for the Detection of Abusive Language (2025.findings-emnlp)

Copied to clipboard

Challenge: Abusive language online creates toxic environments and exacerbates social tensions, underscoring the need for robust NLP models to interpret nuanced linguistic cues.
Approach: They propose a Gated Attention mechanism that combines the strengths of Contextual attention and Self-attention mechanisms to address the limitations of existing attention models within the text classification task.
Outcome: The novel gated attention mechanism addresses the limitations of existing attention models within the text classification task.
Safe-FedLLM: Delving into the Safety of Federated Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing work on federated learning for large language models (FL) addresses privacy and data-silo issues in the training of large language model training.
Approach: They propose a probe-based defense framework for FedLLM that constructs defenses across three levels: Step-Level, Client-Level and Shadow-Level.
Outcome: The proposed framework improves FedLLM's robustness against malicious clients while maintaining competitive performance on benign data.
PERSEVAL: A Framework for Perspectivist Classification Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Perspectivist evaluation practices in NLP remain fragmented and inconsistent .
Approach: They propose a framework that evaluates perspectivist models at the individual annotator level and treats annotators and users as distinct entities, consistent with real-world scenarios.
Outcome: The proposed framework evaluates annotators and users as distinct entities consistent with real-world scenarios.
Linear Layer Extrapolation for Fine-Grained Emotion Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that Transformer-based language models are more factual accurate in later layers .
Approach: They propose a method that optimizes contrast based on the selected intermediate layer . they observe a similar pattern for fine-grained emotion classification in text .
Outcome: Experiments show that the proposed method outperforms standard methods in fine-grained emotion classification tasks.
Recognizing Social Cues in Crisis Situations (2024.lrec-main)

Copied to clipboard

Challenge: During natural disasters, observations of other people's behavior can play an essential role in a person's decision-making.
Approach: They propose a task to categorize social cues in tweets during crisis situations using an annotated dataset of 6,000 tweets.
Outcome: The proposed task is challenging for existing systems and a manual task is based on a dataset of 6,000 tweets labeled with eight social cue categories.
Reference-guided Style-Consistent Content Transfer (2024.lrec-main)

Copied to clipboard

Challenge: Text style transfer involves changing the style of a text while preserving its original style.
Approach: They propose a task of style-consistent content transfer which involves modifying a text’s content based on a provided reference statement while preserving its original style.
Outcome: The proposed approach meets three important conditions: reference faithfulness, style adherence, and coherence.
Mitigating Spurious Correlations in Text Classification Using Latent Space Geometry (2026.acl-long)

Copied to clipboard

Challenge: Existing models rely on predictive shortcuts that hold in training data but break under distribution shifts, leading to large performance drops for minority groups.
Approach: They propose a framework that transforms abstract biases into interpretable geometric anchors without auxiliary classifiers by manipulating latent space geometry.
Outcome: The proposed framework outperforms state-of-the-art baselines and improves worst-group accuracy by over 20% on the CivilComments dataset.
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis.
Approach: They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them.
Outcome: The proposed package is the first to be released in Bulgarian and is available for free.
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data (2026.acl-long)

Copied to clipboard

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, null Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
Challenge: Language identification (LID) is a fundamental step in curating multilingual corpora.
Approach: They introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages.
Outcome: The proposed benchmark covers 109 languages and shows that existing evaluations overestimate accuracy for many languages in the web domain.
Synergizing Semantic Anchors and Ordinal Smoothed Cross-Entropy for Speech Fluency Classification (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods fail to bridge the semantic gap between static expert priors and dynamic temporal representations while overlooking the inherent ordinal nature of fluency scores.
Approach: They propose a set of expert features targeting fluency disruptions and rhythmic regularity to provide explicit linguistic priors.
Outcome: The proposed model outperforms baseline models in both macroscopic and microscopic speech flow trends and local anomalies.
MoVa: Towards Generalizable Classification of Human Morals and Values (2025.emnlp-main)

Copied to clipboard

Challenge: Identifying human morals and values embedded in language is essential to empirical studies of communication.
Approach: They propose a framework for generalizable classification of human morals and values . they recommend a classification strategy that scores all related concepts simultaneously .
Outcome: The proposed method outperforms fine-tuned models across domains and frameworks.
N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal Vocalizations (2025.emnlp-main)

Copied to clipboard

Challenge: Nonverbal vocalizations are an essential component of human communication, conveying rich information without linguistic content.
Approach: They propose a backbone-agnostic framework to disentangle emotion and speaker information from nonverbal vocalizations by leveraging N views of audio samples to learn invariance to specific transformations.
Outcome: The proposed framework achieves competitive performance compared to state-of-the-art methods on the VIVAE, ReCANVo, and ReCANVO-Balanced datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations